Skip to main navigation Skip to search Skip to main content

Adjective density as a text formality characteristic for automatic text classification: A study based on the British national corpus

    Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

    Abstract

    In this article, we report significant findings resulting from an investigation into the correlation between adjective density, calculated as the proportion of adjectives in word tokens, and degrees of text formality as part of an attempt to examine the potential application of adjectives in automatic text classification and identification. Correlations obtained from the training corpus will be compared with human ranking of the text categories concerned in the study and then adapted to unseen data in the test set. A linear regression analysis suggests a strong correlation between degrees of text formality and adjective density. With a weighted average F-measure of 0.606 achieved by a Naïve Bayes classifier, the research establishes adjectives as a powerful differentia of text categories amongst the open word classes, an important feature that has been generally ignored by past studies in automatic text categorization. The empirical findings suggest that the use of adjective density will lead to enhanced practical systems for automatic text classification. © 2009 by Alex Chengyu Fang and Jing Cao.
    Original languageEnglish
    Title of host publicationPACLIC 23 : Proceedings of the 23rd Pacific Asia Conference on Language, Information and Computation
    EditorsOlivia Kwong
    PublisherCity University of Hong Kong Press
    Pages130-139
    Volume1
    ISBN (Print)9789624423198
    Publication statusPublished - Dec 2009
    Event23rd Pacific Asia Conference on Language, Information and Computation (PACLIC 23) - Prince Restaurant, Hong Kong, China
    Duration: 3 Dec 20095 Dec 2009
    http://paclic23.ctl.cityu.edu.hk/PACLIC23_venue.html

    Conference

    Conference23rd Pacific Asia Conference on Language, Information and Computation (PACLIC 23)
    Abbreviated titlePACLIC 23
    PlaceChina
    CityHong Kong
    Period3/12/095/12/09
    Internet address

    Bibliographical note

    Full text of this publication does not contain sufficient affiliation information. The Research Unit(s) information for this record is based on the corrigendum and the then academic department affiliation of the author(s).

    Research Keywords

    • Adjective density
    • Linear regression
    • Naïve bayes
    • Text classification
    • Text formality

    Fingerprint

    Dive into the research topics of 'Adjective density as a text formality characteristic for automatic text classification: A study based on the British national corpus'. Together they form a unique fingerprint.

    Cite this