Abstract
Taxonomies of the Web typically have hundreds of thousands of categories and skewed category distribution over documents. It is not clear whether existing text classification technologies can perform well on and scale up to such large-scale applications. To understand this, we conducted the evaluation of several representative methods (Support Vector Machines, k-Nearest Neighbor and Naive Bayes) with Yahoo! taxonomies. In particular, we evaluated the effectiveness/efficiency tradeoff in classifiers with hierarchical setting compared to conventional (flat) setting, and tested popular threshold tuning strategies for their scalability and accuracy in large-scale classification problems.
Copyright is held by the author/owner(s).
Copyright is held by the author/owner(s).
| Original language | English |
|---|---|
| Title of host publication | 14th International World Wide Web Conference, WWW2005 |
| Pages | 1106-1107 |
| DOIs | |
| Publication status | Published - 2005 |
| Externally published | Yes |
| Event | 14th International World Wide Web Conference, WWW2005 - Chiba, Japan Duration: 10 May 2005 → 14 May 2005 |
Publication series
| Name | 14th International World Wide Web Conference, WWW2005 |
|---|
Conference
| Conference | 14th International World Wide Web Conference, WWW2005 |
|---|---|
| Place | Japan |
| City | Chiba |
| Period | 10/05/05 → 14/05/05 |
Bibliographical note
Publication details (e.g. title, author(s), publication statuses and dates) are captured on an “AS IS” and “AS AVAILABLE” basis at the time of record harvesting from the data source. Suggestions for further amendments or supplementary information can be sent to [email protected].Research Keywords
- Algorithm complexity
- Parameter tuning strategies
- Text categorization
- Very large web taxonomies
Fingerprint
Dive into the research topics of 'An experimental study on large-scale web categorization'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver