Skip to main navigation Skip to search Skip to main content

Word segmentation in Chinese language processing

  • Xinxin Shu
  • , Junhui Wang
  • , Xiaotong Shen
  • , Annie Qu*
  • *Corresponding author for this work

Research output: Journal Publications and ReviewsRGC 21 - Publication in refereed journalpeer-review

Abstract

This paper proposes a new statistical learning method for word segmentation in Chinese language processing. Word segmentation is the crucial first step towards natural language processing. Segmentation, despite progress, remains under-studied; particularly for the Chinese language, the second most popular language among all internet users. One major difficulty is that the Chinese language is highly context-dependent and ambiguous in terms of word representations. To overcome this difficulty, we cast the problem of segmentation into a framework of sequence classification, where an instance (observation) is a sequence of characters, and a class label is a sequence determining how each character is segmented. Given the class label, each character sequence can be segmented into linguistically meaningful words. The proposed method is investigated through the Peking university corpus of Chinese documents. Our numerical study shows that the proposed method compares favorably with the state-of-the-art segmentation methods in the literature.
Original languageEnglish
Pages (from-to)165-173
JournalStatistics and Its Interface
Volume10
Issue number2
Online published31 Oct 2016
Publication statusPublished - Jan 2017

Funding

Research supported in part by National Science Foundation Grants DMS-1207771, DMS-1415500, DMS-1415308, DMS-1308227, DMS-1415482, and HK GRF-11302615. The authors thank the editors, the associate editor and the reviewers for helpful comments and suggestions.

Research Keywords

  • Cutting-plane algorithm
  • Language processing
  • Support vector machines
  • Word segmentation
  • PROBABILITY ESTIMATION
  • SELECTION
  • LIKELIHOOD

Fingerprint

Dive into the research topics of 'Word segmentation in Chinese language processing'. Together they form a unique fingerprint.

Cite this