Skip to main navigation Skip to search Skip to main content

A new suffix tree similarity measure for document clustering

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

Abstract

In this paper, we propose a new similarity measure to compute the pairwise similarity of text-based documents based on suffix tree document model. By applying the new suffix tree similarity measure in Group-average Agglomerative Hierarchical Clustering (GAHC) algorithm, we developed a new suffix tree document clustering algorithm (NSTC). Experimental results on two standard document clustering benchmark corpus OHSUMED and RCV1 indicate that the new clustering algorithm is a very effective document clustering algorithm. Comparing with the results of traditional word term weight tf-idf similarity measure in the same GAHC algorithm, NSTC achieved an improvement of 51% on the average of F-measure score. Furthermore, we apply the new clustering algorithm in analyzing the Web documents in online forum communities. A topic oriented clustering algorithm is developed to help people in assessing, classifying and searching the the Web documents in a large forum community.
Copyright is held by the International World Wide Web Conference Committee (IW3C2).
Original languageEnglish
Title of host publication16th International World Wide Web Conference, WWW2007
Pages121-130
DOIs
Publication statusPublished - 2007
Event16th International World Wide Web Conference, WWW2007 - Banff, AB, Canada
Duration: 8 May 200712 May 2007

Publication series

Name16th International World Wide Web Conference, WWW2007

Conference

Conference16th International World Wide Web Conference, WWW2007
PlaceCanada
CityBanff, AB
Period8/05/0712/05/07

Bibliographical note

Publication details (e.g. title, author(s), publication statuses and dates) are captured on an “AS IS” and “AS AVAILABLE” basis at the time of record harvesting from the data source. Suggestions for further amendments or supplementary information can be sent to [email protected].

Funding

The research work described in this paper is partially supported by a City University TDF grant [Project No. 6980080] and a SRG grant of City University of Hong Kong [Project No. 7001975].

Research Keywords

  • Document model
  • Similarity measure
  • Suffix tree

Fingerprint

Dive into the research topics of 'A new suffix tree similarity measure for document clustering'. Together they form a unique fingerprint.

Cite this