Abstract
Automatic Term Extraction (ATE) has gained increasing prominence in natural language processing with the development of electronic corpora. Many efforts have been made to improve ATE processing by means of different kinds of techniques. Actual terminological practice has already shown that context is an important factor for term identification. Considering verbs as such an important syntactic factor, this paper attempts to explore the contribution that verb syntax information will offer to ATE. This paper will aim at measuring probabilistic associations between terms and syntactic constituents within the syntactic structure of the sentence. More specifically, this study adopts the classification of verb patterns provided in Oxford Advanced Learner's Dictionary (6th edition), which has 22 different verb patterns. 100,000 abstracts from MEDLINE* with the keyword ‘internal medicine’ are collected as a primary corpus. Four sub-corpora were created, each with 2,000 randomized abstracts from the primary corpus. The four testing corpora, totaling 400,883 words in 15,880 sentences, are subsequently parsed by the Stanford Parser. A list of medical terms was created from Medical Subject Headings, a controlled vocabulary thesaurus from National Library of Medicine. All of the terms in the parse trees are annotated according to those in the list. The result turns out to be quite indicative in terms of predicting the syntactic positions with higher likelihood for term occurrences. Across the four sub-corpora, the verb patterns with the highest number of term occurrences are [VN] **, [VN+adv./prep.] **and [V-N]**. Moreover, even though the rest of the patterns are not ranked exactly in the same order across the four sub-corpora, we can still have better understanding if we group the minor patterns together. For example, if we consider the three verb patterns, [V-ADJ] **, [V+adv./prep.] ** and [VNN] * *as one group, it can be seen that this group is ranked as the fourth highest across the four sub-corpora. As a further example, if [V+(that)]**, [V+to inf]** and [VN+ to inf] **are grouped together, it will be ranked as the fifth across the four sub-corpora. As a result, the conclusion can be drawn that term occurrences are quite consistent across the four sub-corpora. Therefore, after training on a large medical corpus drawn from the MEDLINE, a model can be created that predicts the likelihood of a noun phrase being a term given a syntactic constituent. When it comes to actual ATE application, higher weights will be given to those term candidates located in these verb patterns for better probabilistic estimates. The study reported in this paper suggests that verb pattern information can be used as a reliable indicator for the calculation of the termhood of term candidates. Moreover, during the linguistic processing of these corpora, it is found that different parsers present linguistic information of different granularities, which will subsequently dictate the accuracy of term recognition. In a future experiment, the exact syntactic positions will be investigated for the three verb patterns with the highest term occurrences. A different parser will be used which produces more detailed syntactic analysis. A focus of the future study will be to investigate to what a extent a different parser will influence the performance of ATE. * MEDLINE (Medical Literature Analysis and Retrieval System Online) is the U.S. National Library of Medicine's (NLM) bibliographic database.** [VN] means a transitive verb with a noun, pronoun or noun phrase as object;[VN+adv./prep.] means transitive words used with a prepositional phrase or an adverb;[V-N] means linking verbs with a noun phrase as the complement;[V-ADJ] means linking verbs that have an adjective as the complement;[V+adv./prep.] means intransitive verbs used with a prepositional phrase or an adverb;[VNN] means verbs used with two objects;[V+(that)]means a verb is followed by a clause beginning with that;[V+to inf] means verbs used with to-infinitive;[VN+ to inf] means verbs used with both a noun phrase and a to-infinitive.
| Original language | English |
|---|---|
| Publication status | Published - 27 May 2009 |
| Event | 30th Annual Conference of the International Computer Archive for Modern and Medieval English - , United Kingdom Duration: 27 May 2009 → 31 May 2009 |
Conference
| Conference | 30th Annual Conference of the International Computer Archive for Modern and Medieval English |
|---|---|
| Place | United Kingdom |
| Period | 27/05/09 → 31/05/09 |
Fingerprint
Dive into the research topics of 'Measuring Probabilistic Relations between Terms and Syntactic Units'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver