Skip to main navigation Skip to search Skip to main content

Building and Analysing Corpora of Computer-Mediated Communication

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 12 - Chapter in an edited book (Author)

Abstract

This chapter addresses problems encountered during the construction and analysis of a synchronic corpus of computer-mediated discourse. The corpus was not primarily constructed for the examination of the linguistic idiosyncrasies of the online chatting medium; rather it is to be used for corpus-based sociolinguistic inquiry into the language use and identity construction of a particular social group (which in this case could be classed as ‘vulnerable’). Therefore the corpus data needed considerable adaptation during compilation and analysis to prevent those idiosyncrasies from acting as noise in the data. Adaptations include responses to spam (in the form of ‘adbots’), cyber-orthography, the ubiquity of names, overlapping conversations and challenges of annotation. Difficulties with gaining participant permissions and demographic information also required significant attention. Attempted solutions to these corpus construction and analysis challenges, which are closely bound to the fields of both cyber-research and corpus linguistics, are outlined.
Original languageEnglish
Title of host publicationContemporary Corpus Linguistics
EditorsPaul Baker
Place of PublicationLondon
PublisherContinuum
Pages301-320
ISBN (Print)9780826496102
Publication statusPublished - 5 Jun 2009
Externally publishedYes

Publication series

NameContemporary Studies in Linguistics
PublisherContinuum

Fingerprint

Dive into the research topics of 'Building and Analysing Corpora of Computer-Mediated Communication'. Together they form a unique fingerprint.

Cite this