- 영문명
- Web Page Similarity based on Size and Frequency of Tokens
- 발행기관
- 한국IT서비스학회
- 저자명
- 이은주(Eun-Joo Lee) 정우성(Woo-Sung Jung)
- 간행물 정보
- 『한국IT서비스학회지』한국IT서비스학회지 제11권 제4호, 263~275쪽, 전체 13쪽
- 주제분류
- 경제경영 > 경영학
- 파일형태
- 발행일자
- 2012.12.31
국문 초록
영문 초록
It is becoming hard to maintain web applications because of high complexity and duplication of web pages. However, most of research about code clone is focusing on code hunks. and their target is limited to a specific language. Thus, we propose GSIM, a language-independent statistical approach to detect similar pages based on scarcity and frequency of customized tokens. The tokens, which can be obtained from pages splitted by a set of given separators, are defined as atomic elements lor calculating similarity between two pages. In this paper, the domain definition for web applications and algorithms for collecting tokens, making matrics, calculating similarity are given. We also conducted experiments on open source codes for evaluation, with our GSIM tool. The results show the applicability of the proposed method and the effects of parameters such as threshold, toughness, length of tokens, on their quality and performance.
목차
Abstract
1. 서론
2. 관련 연구
3. GSIM 정의
4. 실험 결과
5. 결론
참고문헌
저자소개
해당간행물 수록 논문
참고문헌
최근 이용한 논문
교보eBook 첫 방문을 환영 합니다!
신규가입 혜택 지급이 완료 되었습니다.
바로 사용 가능한 교보e캐시 1,000원 (유효기간 7일)
지금 바로 교보eBook의 다양한 콘텐츠를 이용해 보세요!