Skip to main navigation Skip to search Skip to main content

Detection of non-native sentences using machine-translated training data

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

Abstract

Training statistical models to detect nonnative sentences requires a large corpus of non-native writing samples, which is often not readily available. This paper examines the extent to which machinetranslated (MT) sentences can substitute as training data. Two tasks are examined. For the native vs non-native classification task, nonnative training data yields better performance; for the ranking task, however, models trained with a large, publicly available set of MT data perform as well as those trained with non-native data.
Original languageEnglish
Title of host publicationNAACL-HLT 2007 - Human Language Technology Conference of the North American Chapter of the Association of Computational Linguistics, Companion Volume: Short Papers
PublisherAssociation for Computational Linguistics
Pages93-96
DOIs
Publication statusPublished - 2007
Externally publishedYes
Event2007 Human Language Technology Conference of the North American Chapter of the Association of Computational Linguistics, NAACL-HLT 2007 - Rochester, United States
Duration: 22 Apr 200727 Apr 2007
https://aclanthology.org/N07-2

Publication series

NameNAACL-HLT 2007 - Human Language Technology Conference of the North American Chapter of the Association of Computational Linguistics, Companion Volume: Short Papers

Conference

Conference2007 Human Language Technology Conference of the North American Chapter of the Association of Computational Linguistics, NAACL-HLT 2007
PlaceUnited States
CityRochester
Period22/04/0727/04/07
Internet address

Bibliographical note

Publication details (e.g. title, author(s), publication statuses and dates) are captured on an “AS IS” and “AS AVAILABLE” basis at the time of record harvesting from the data source. Suggestions for further amendments or supplementary information can be sent to [email protected].

Fingerprint

Dive into the research topics of 'Detection of non-native sentences using machine-translated training data'. Together they form a unique fingerprint.

Cite this