Abstract
We investigate the possibility of exploiting character-based dependency for Chinese information processing. As Chinese text is made up of character sequences rather than word sequences, word in Chinese is not so natural a concept as in English, nor is word easy to be defined without argument for such a language. Therefore we propose a character-level dependency scheme to represent primary linguistic relationships within a Chinese sentence. The usefulness of character dependencies are verified through two specialized dependency parsing tasks. The first is to handle trivial character dependencies that are equally transformed from traditional word boundaries. The second furthermore considers the case that annotated internal character dependencies inside a word are involved. Both of these results from character-level dependency parsing are positive. This study provides an alternative way to formularize basic character-and word-level representation for Chinese. © 2009 Association for Computational Linguistics.
| Original language | English |
|---|---|
| Title of host publication | EACL 2009 - 12th Conference of the European Chapter of the Association for Computational Linguistics, Proceedings |
| Publisher | ACL Anthology |
| Pages | 879-887 |
| ISBN (Print) | 9781932432169 |
| DOIs | |
| Publication status | Published - 2009 |
| Event | 12th Conference of the European Chapter of the Association for Computational Linguistics , EACL 2009 Student Research Workshop - Athens, Greece Duration: 30 Mar 2009 → 3 Apr 2009 |
Publication series
| Name | EACL 2009 - 12th Conference of the European Chapter of the Association for Computational Linguistics, Proceedings |
|---|
Conference
| Conference | 12th Conference of the European Chapter of the Association for Computational Linguistics , EACL 2009 Student Research Workshop |
|---|---|
| Place | Greece |
| City | Athens |
| Period | 30/03/09 → 3/04/09 |
Bibliographical note
Publication details (e.g. title, author(s), publication statuses and dates) are captured on an “AS IS” and “AS AVAILABLE” basis at the time of record harvesting from the data source. Suggestions for further amendments or supplementary information can be sent to [email protected].Funding
This work is beneficial from many sources, including three anonymous reviewers. Especially, the authors are grateful to two colleagues, one reviewer from EMNLP-2008 who gave some very insightful comments to help us extend this work, and Mr. SONG Yan who annotated internal dependencies of top frequent 22K words extracted from UPUC segmentation corpus. Of course, it is the duty of the first author if there still exists anything wrong in this work.
Fingerprint
Dive into the research topics of 'Character-level dependencies in Chinese: Usefulness and learning'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver