Skip to main navigation Skip to search Skip to main content

Develop a Large Scale Chinese Reading Corpus for Machine Learning of Unconscious Word Segmentation from Eye Movements

Project: Research

Project Details

Description

(1) Large scale reading corpora recording readers’ eye movements during reading haverarely been used to study Chinese reading, where two key questions remain unanswered:(2) how words are segmented during reading Chinese texts (without spaces as visuallysalient word delimiters) and (3) whether eye movements hint an answer to (2); and (4)wordhood remains a notorious long-standing issue unlikely to be resolved by worddefinition, theoretical argumentation and/or technological advancement in both Chineselinguistics and language processing. Given this background, this project is proposed todevelop a large scale corpus of Chinese reading for machine learning to learn from eyemovements how native readers segment words. Regarding inconsistent results in Chinesereading research due, to a large extent, to typographic settings (e.g., adequate inter- vs.intra-word spacing), we will first identify their optimal ranges, to ensure the validity,reliability and consistency throughout our reading experiments. Then, we will conduct apilot study to (re-)examine word spacing effect with spaces/delimiters of various sizes,shapes and colors, to demonstrate the significance of optimal typographic settings,before moving on to developing the reading corpus by recording readers’ eye movementsduring reading under such settings. The reading materials, presented to readers withoutword boundary information, are selected from representative segmented Chinese corporathat conform to a Chinese word segmentation standard. The key task to enable themachine learning is the feature engineering to convert eye movements to an optimal setof features (within a predefined feature space defined on characters within theperceptual span of reading with respect to available variables in eye tracking reports)that give the best model for predicting word boundaries after training. Our previousworks on machine learning of Chinese word segmentation via character tagging usingthe conditional random fields model and on large scale feature selection for semanticparsing, that advanced the state of the art in respective areas, provide a strong technicalsupport and plenty practical experience to propel our success in this learning task. Asthe very first endeavor on (1) in the field, this project will produce not only reliablebenchmark data for Chinese reading but also verifiable answers to (2) and (3), by meansof modeling how readers’ eye movements reveal their unconscious (vs. subjective) wordsegmentation, and an objective means to scrutinize how native speakers processcontroversial cases of Chinese wordhood, for novel insights into this perplexing issue.
Project number9042276
Grant typeGRF
StatusFinished
Effective start/end date1/01/1626/06/20

Keywords

  • Chinese reading ,Eye tracking,Eye movement,Typographic feature,

Fingerprint

Explore the research topics touched on by this project. These labels are generated based on the underlying awards/grants. Together they form a unique fingerprint.