Skip to main navigation Skip to search Skip to main content

Scalable Authorship Identification

Student thesis: Doctoral Thesis

Abstract

The science of authorship identification is based on the observation that there exists unconscious elements of literary styles that can help differentiate among the writing styles of the authors. Authorship identification aims at identifying the true author of an anonymous document from a set of candidate authors. Authorship identification problem has several variations such as (i) authorship identification of single-author documents (AISD); (ii) authorship identification of multi-author documents (AIMD); and (iii) crosslingual authorship identification (CLAI).

Authorship identification of single-author documents (AISD) aims at identifying the true author of an anonymous single-author document from a set of candidate authors. Recently, the practical applications of AISD have grown in several areas such as intelligence agencies work, e.g., linking intercepted messages to known enemies or terrorists; criminal law, e.g., identifying the writers of harassing letters or ransom notes; plagiarism detection, e.g., determining whether work submitted by a student was written by someone else. An application of authorship attribution can also be found in the area of digital humanities, where issues of interest include authentication of disputed literary text. Meanwhile, due to the availability of large text repositories on the Internet, the problem of managing them becoming ever more important. For example, categorizing text documents by their authors has been receiving increasing research attention in the areas of web information management, information retrieval, and statistical natural language processing . Existing AISD studies report a high accuracy (over 90%) using several types of stylometric features including lexical, character, syntactic and structural ones. However, these studies have the following limitations which makes them inapplicable to real world problems such as plagiarism detection. Firstly, existing AISD studies report a drastic drop in the accuracy as the number of candidate authors increases. In fact, in most existing AISD studies reporting a high accuracy (90% or over), the number of candidate authors does not exceed 50. However, the real world applications tend to have hundreds of authors. As a result, building a predictive model that can accurately identify the author of an anonymous text is a challenging task. Secondly, a slight variation in the document length affects the accuracy of AISD. This is not a problem in a controlled scenario where researchers can manually choose a small subset of documents with similar lengths to study. However, in a real world dataset (such as a database of student essays from different courses), there can be a substantial variation among the lengths of the available texts.

Authorship identification of multi-author documents (AIMD) aims at identifying the true authors of a multi-author document. AIMD can be formally defined as follows. Given a corpus of multi-author documents labeled with their co-authors, identify the co-authors of an anonymous multi-author document from the authors in the given corpus. One prominent application domain of AIMD is bibliometrics, in which AIMD can help improve the processes of measuring and analyzing the collaborative natures among a community of researchers. Instead of attributing the entire paper to all the listed authors, one can use AIMD techniques to perform a more fine-grained analysis. Specifically, different parts of the same document can be attributed to different authors on the author list. Such an authorship identification capability can help the information retrieval system in the following ways: (i) scholarly search engines may implement an author specific search in which the researchers can look for text sample written by a particular author; and (ii) a researcher may wish to construct individual author profiles reflecting the contributions of each author in different scientific fields. In addition, AIMD techniques can also be used to identify researchers who had been involved actively in writing and mentors who are giving feedback and providing ideas. Another aspect of the AIMD is the peer-review system of the academic conferences where both the reviewers and the authors of the paper stay anonymous. This notion can be challenged by showing that it is possible for a reviewer to reveal the identity of the authors of scientific papers by using the AIMD framework.

Existing authorship identification techniques designed to handle single-author documents are not applicable to multi-author documents. This is because, single-author authorship attribution techniques rely on the assumption that every text sample (document) has only one single label (author). However, the AIMD problem requires the ability to (i) infer the writing style of each individual author from a corpus of multi-author documents; and (ii) make a multi-label prediction for each document.

Existing AIMD studies have the following limitations. (i) The accuracy levels of existing AIMD techniques can still be greatly improved. For example, the state-of-the-art stylometry based technique reports an accuracy level less than 30% on a corpus containing over 360 candidate authors. (ii) Existing techniques are adversely affected by an increase in the number of co-authors. For example, the best existing AIMD method reported a drop in accuracy level from 25% to 16% as the number of authors had increased from 2 to 7. (iii) Existing AIMD techniques do not tackle the issue of non-writing authors (NWA) [5, 16]. However, NWAs do exist in real world scenarios. For example, in a scientific/engineering article, it is not necessary that all listed co-authors had contributed as writers.

Cross-lingual authorship identification (CLAI) aims at finding the author of an anonymous document written in one language by using labeled documents written in other languages. Most of the existing research in this area has used mono-lingual corpora and English is the most studied language. However, nowadays, users may participate in several platforms regardless of the language. For example, an Italian user may have a blog in Italian, primarily post in English on Facebook, and publish articles in both languages. Similarly, many novelists write in different languages, for example, Vladimir Nabokov wrote in both Russian and English and an Irish novelist Samuel Beckett wrote in both French and English. Moreover, nowadays, around 45% content on the web is written in non-English languages. Another aspect is that people are becoming increasingly proficient in more than one language. It has been shown that more than half of world population is bilingual. The European Union report shows that on average 94.5% pupils in secondary education learn two or more languages. Consequently, there is a substantial need for cross-lingual authorship identification solutions.

The main challenge of cross-lingual authorship identification is that stylistic markers (features) used in one language may not be applicable to other languages in the corpus. Existing methods overcome this challenge by using external resources such as machine translation and part-of-speech tagging. However, such solutions are not applicable to languages with poor external resources (known as low resource languages). Moreover, these existing methods also fail to scale as the number of candidate authors and/or the number of languages in the corpus increase.

The main contributions of this thesis for solving the aforementioned challenges of each problem are outlined as follows:

In order to address the limitations of the AISD problem, we propose a novel stylometric data representation model which can handle (i) documents with different lengths; and (ii) a large number of candidate authors. Specifically, we represent each document as a set of data points in a high dimensional space. In this set representation, each data point represents a chunk (a document segment with a fixed size) and each dimension corresponds to a stylometric feature.

In order to address the limitations of the AIMD problem, we propose an AIMD technique called Co-Authorship Graph (CAG) which can be used to collaboratively attribute text fragments in a set of documents to distinguish authors in the same community. Based on the CAG technique, we propose a novel AIMD solution which (i) significantly outperforms the existing state-of-the-art solution; (ii) can effectively handle a larger number of co-authors; and (iii) is capable of handling non-writing authors (NWA).

In order to address the limitations of CLAI, we identify a set of features that can be used across a large number of languages. Specifically, our feature space relies on a minimal set of linguistic assumptions: (i) the ability to tokenize a writing sample into words; (ii) the ability to identify sentence boundaries; and (iii) the use of punctuations. Our proposed cross-lingual authorship identification (CLAI) method does not rely on machine translation or any internal knowledge of the languages in the corpus.
Date of Award20 Aug 2018
Original languageEnglish
Awarding Institution
  • City University of Hong Kong
SupervisorSarana NUTANONG (Supervisor)

Cite this

'