The field of sentiment analysis and opinion mining, which is concerned with automatically analyzing sentiments and opinions in textual sources (e.g., news, blogs, reviews and tweets), has attracted much attention from researchers and practitioners in recent years. Its potential applications include review summarization and product recommendation, opinion retrieval, political polling, and sentiment-aware online advertising.
Much previous work on sentiment analysis in the Natural Language Processing (NLP) field has relied on supervised machine learning approaches, which use labeled data to learn computational models for sentiment analysis. However, labeled data are often difficult, expensive, or time consuming to obtain because the efforts of experienced human annotators may be required. Moreover, sentiment analysis is a problem with multi-facets (e.g., subjectivity, polarity, opinion holder/target, etc.) and all of them suffer from this lack of labeled data.
This study aims to explore rich linguistic knowledge and unlabeled data to build better computational models for sentiment analysis with a view to reducing the need of more labeled data for performance improvement. Linguistically motivated approaches and semi-supervised learning could respectively address this lack of labeled data by using rich linguistic knowledge and large amounts of unlabeled data, together with labeled data, to build better computational models for sentiment analysis, and the two points are equally applicable to the multi-facets of sentiment analysis. This thesis represents substantially novel research which is significant in that a series of linguistically motivated and semi-supervised approaches are presented for addressing various demands of sentiment analysis from different facets:
First, for word-level sentiment analysis, we propose a linguistically motivated and semi-supervised approach to learning Chinese polarity lexicons by using both synonym relations among words and internal morphological features within Chinese words (i.e., by integrating two types of different but complementary models: synonym-based graph models and morphological feature-based models).
Second, for sentence-level subjectivity classification, we investigate lexical item-based supervised approaches to opinionated sentence recognition by using ensemble techniques. Because sentiment lexicons and labeled corpora both represent human understanding of sentiment, the integration of them under ensemble techniques can make full use of these two types of resources.
Third, we propose to identify opinion holders/targets in news sentences with syntactic dependency structure, i.e., to identify opinion holders by means of reporting verbs and to identify opinion targets by considering both opinion holders and sentiment-bearing words.
Fourth, we present an innovative model for multilingual sentiment classification based on parallel information, which augments available labeled data in each language with unlabeled parallel data. The proposed model jointly learns improved monolingual sentiment classifiers at the sentence level for each language, based on the intuition that the sentiment labels for parallel sentences should be similar.
Last, based on a large amount of unlabeled data, we investigate the role of semantic information obtained with weakly supervised topic models to two subproblems in multi-aspect sentiment analysis, i.e., sentence-level aspect classification and multi-aspect sentiment rating prediction, and then show their application on multi-aspect review summarization.
The above contributions are of value to the development of pratical sentiment analysis systems, and also allow for a deeper understanding of sentiments from the linguistic and computational perspectives.
| Date of Award | 15 Jul 2013 |
|---|
| Original language | English |
|---|
| Awarding Institution | - City University of Hong Kong
|
|---|
| Supervisor | Ka Yin Benjamin T'SOU (Supervisor) |
|---|
- Supervised learning (Machine learning)
- Computational linguistics
- Content analysis (Communication)
On computing textual sentiment with linguistic knowledge and semi-supervised learning
LU, B. (Author). 15 Jul 2013
Student thesis: Doctoral Thesis