Abstract
We used a production segmentation system, which draws heavily on a large dictionary derived from processing a large amount (over 150 million Chinese characters) of synchronous textual data gathered from various Chinese speech communities, including Beijing, Hong Kong, Taipei, and others. We run this system in two tracks in the Second International Chinese Word Segmentation Bakeoff, with Backward Maximal Matching (right-to-left) as the primary mechanism. We also explored the use of a number of supplementary features offered by the large dictionary in post-processing, in an attempt to resolve ambiguities and detect unknown words. While the results might not have reached their fullest potential, they nevertheless reinforced the importance and usefulness of a large dictionary as a basis for segmentation, and the implication of following a uniform standard on the segmentation performance on data from various sources. © 2005 SIGHAN@IJCNLP 2005 - 4th SIGHAN Workshop on Chinese Language Processing, Proceedings of the Workshop. All Rights Reserved.
| Original language | English |
|---|---|
| Title of host publication | SIGHAN@IJCNLP 2005 - 4th SIGHAN Workshop on Chinese Language Processing, Proceedings of the Workshop |
| Publisher | ACL Anthology |
| Pages | 176-179 |
| Publication status | Published - 2005 |
| Event | 4th SIGHAN Workshop on Chinese Language Processing at the 2nd International Joint Conference on Natural Language Processing, SIGHAN@IJCNLP 2005 - Jeju Island, Korea, Republic of Duration: 14 Oct 2005 → 15 Oct 2005 https://aclanthology.org/volumes/I05-3/ |
Publication series
| Name | SIGHAN |
|---|---|
| Publisher | IJCNLP 2005 - 4th SIGHAN Workshop on Chinese Language Processing, Proceedings of the Workshop |
Conference
| Conference | 4th SIGHAN Workshop on Chinese Language Processing at the 2nd International Joint Conference on Natural Language Processing, SIGHAN@IJCNLP 2005 |
|---|---|
| Place | Korea, Republic of |
| City | Jeju Island |
| Period | 14/10/05 → 15/10/05 |
| Internet address |
Bibliographical note
Publication details (e.g. title, author(s), publication statuses and dates) are captured on an “AS IS” and “AS AVAILABLE” basis at the time of record harvesting from the data source. Suggestions for further amendments or supplementary information can be sent to [email protected].Fingerprint
Dive into the research topics of 'Maximal match chinese segmentation augmented by resources generated from a very large dictionary for post-processing'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver