Abstract
In this paper, we propose a new similarity measure to compute the pairwise similarity of text-based documents based on suffix tree document model. By applying the new suffix tree similarity measure in Group-average Agglomerative Hierarchical Clustering (GAHC) algorithm, we developed a new suffix tree document clustering algorithm (NSTC). Experimental results on two standard document clustering benchmark corpus OHSUMED and RCV1 indicate that the new clustering algorithm is a very effective document clustering algorithm. Comparing with the results of traditional word term weight tf-idf similarity measure in the same GAHC algorithm, NSTC achieved an improvement of 51% on the average of F-measure score. Furthermore, we apply the new clustering algorithm in analyzing the Web documents in online forum communities. A topic oriented clustering algorithm is developed to help people in assessing, classifying and searching the the Web documents in a large forum community.
Copyright is held by the International World Wide Web Conference Committee (IW3C2).
Copyright is held by the International World Wide Web Conference Committee (IW3C2).
| Original language | English |
|---|---|
| Title of host publication | 16th International World Wide Web Conference, WWW2007 |
| Pages | 121-130 |
| DOIs | |
| Publication status | Published - 2007 |
| Event | 16th International World Wide Web Conference, WWW2007 - Banff, AB, Canada Duration: 8 May 2007 → 12 May 2007 |
Publication series
| Name | 16th International World Wide Web Conference, WWW2007 |
|---|
Conference
| Conference | 16th International World Wide Web Conference, WWW2007 |
|---|---|
| Place | Canada |
| City | Banff, AB |
| Period | 8/05/07 → 12/05/07 |
Bibliographical note
Publication details (e.g. title, author(s), publication statuses and dates) are captured on an “AS IS” and “AS AVAILABLE” basis at the time of record harvesting from the data source. Suggestions for further amendments or supplementary information can be sent to [email protected].Funding
The research work described in this paper is partially supported by a City University TDF grant [Project No. 6980080] and a SRG grant of City University of Hong Kong [Project No. 7001975].
Research Keywords
- Document model
- Similarity measure
- Suffix tree
Fingerprint
Dive into the research topics of 'A new suffix tree similarity measure for document clustering'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver