Abstract
In this article, we report significant findings resulting from an investigation into the correlation between adjective density, calculated as the proportion of adjectives in word tokens, and degrees of text formality as part of an attempt to examine the potential application of adjectives in automatic text classification and identification. Correlations obtained from the training corpus will be compared with human ranking of the text categories concerned in the study and then adapted to unseen data in the test set. A linear regression analysis suggests a strong correlation between degrees of text formality and adjective density. With a weighted average F-measure of 0.606 achieved by a Naïve Bayes classifier, the research establishes adjectives as a powerful differentia of text categories amongst the open word classes, an important feature that has been generally ignored by past studies in automatic text categorization. The empirical findings suggest that the use of adjective density will lead to enhanced practical systems for automatic text classification. © 2009 by Alex Chengyu Fang and Jing Cao.
| Original language | English |
|---|---|
| Title of host publication | PACLIC 23 : Proceedings of the 23rd Pacific Asia Conference on Language, Information and Computation |
| Editors | Olivia Kwong |
| Publisher | City University of Hong Kong Press |
| Pages | 130-139 |
| Volume | 1 |
| ISBN (Print) | 9789624423198 |
| Publication status | Published - Dec 2009 |
| Event | 23rd Pacific Asia Conference on Language, Information and Computation (PACLIC 23) - Prince Restaurant, Hong Kong, China Duration: 3 Dec 2009 → 5 Dec 2009 http://paclic23.ctl.cityu.edu.hk/PACLIC23_venue.html |
Conference
| Conference | 23rd Pacific Asia Conference on Language, Information and Computation (PACLIC 23) |
|---|---|
| Abbreviated title | PACLIC 23 |
| Place | China |
| City | Hong Kong |
| Period | 3/12/09 → 5/12/09 |
| Internet address |
Bibliographical note
Full text of this publication does not contain sufficient affiliation information. The Research Unit(s) information for this record is based on the corrigendum and the then academic department affiliation of the author(s).Research Keywords
- Adjective density
- Linear regression
- Naïve bayes
- Text classification
- Text formality
Fingerprint
Dive into the research topics of 'Adjective density as a text formality characteristic for automatic text classification: A study based on the British national corpus'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver