Linguistic activities are basically divided into two groups: production (speaking/writing) and reception (listening/reading). In the context of lexicography, the former is expressed by the term ‘encoding’, and the latter by the term ‘decoding’. This distinction is especially important in the field of bilingual lexicography, because the difference between encoding and decoding processes impacts on the fundamental design of bilingual dictionaries.
There is an urgent and significant need to suggest a better model for encoding bilingual dictionaries. To do this, the actual encoding process must be analysed in detail, breaking free from lexical and part-of-speech boundaries, by comparing different parts of speech, and comparing a word with a unit larger than a single word, between the two languages.
This thesis investigates the process of bilingual encoding from Japanese to English. The significance of this research lies in an attempt to match meanings in the source language not merely by exploring lexical equivalents with the same part of speech in the target language, but instead by expanding the choices to other parts of speech and beyond, to grammatical constructions. These grammatical constructions were analysed within a theoretical framework of a conceptual space and semantic map.
Incongruity between different languages in terms of the parts of speech used to express the same meanings is widely reported in several fields of linguistics. Yet there is an implicit and groundless belief that an idea or concept in one language can or should be expressed in the same part of speech or at the same lexical level in another language. This fact can be evidenced from the obvious design flaws observed in the current bilingual encoding dictionary models. This research suggests and develops a ‘bilingual semantic map model’ for Japanese and English part-of-speech constructions to match meanings between the two languages.
The theoretical backbone of this thesis is construction grammar, in the field of cognitive linguistics. The validity of the theoretical framework, such as conceptual space and semantic map, is proven by the frequent use of the concepts especially by typologists. However, application of these notions is scarce in other fields, for instance, in lexicography or translation studies. The value of part-of-speech information goes
almost unchallenged in those fields.
A corpus-based approach was adopted in this research, using one parallel corpus and two comparable corpora. The analysis was based on 150,000 pairs of morphologically and syntactically annotated sentences in Japanese and English. The methodology was carefully designed to select a parallel corpus in which the same meanings were computationally aligned. This alignment of 150,000 pairs of Japanese and English sentences enables us to concentrate our focus on the analysis of constructions. Given that the semantics between the two languages is fixed, the use of similar or different parts of speech and constructions in the corpus reveal the strategies employed in the process of bilingual encoding from Japanese to English.
The results show that even the content words, i.e. nouns, verbs, adjectives, and adverbs, exhibit visible distributional differences between the two languages. The proportion of nouns observed in the parallel corpus in Japanese and English is 58.09% and 41.91% respectively; Japanese uses around 16% more nouns than English. For verbs, the ratio reverses. English uses more verbs than Japanese. For Japanese, the proportion is 42.76% and for English, 57.24%. The differences are more striking in the comparison of adjectives, and also of adverbs. The proportion of adjectives is 25.65% for Japanese and 74.35% for English. Similarly, the proportion of adverbs is 26.90% for Japanese and 73.10% for English. These huge discrepancies are proof that parts of speech cannot be a key factor in navigating the bilingual encoding processes. Based on these findings, the research groups and compares constructions from each language, and places them in a bilingual semantic map that includes frequency information. This map reveals how the two languages are connected.
The results of this research demonstrate that lexis and constructions cannot be separated but rather that they are closely interwoven in instantiating meanings in the process of bilingual encoding. This finding has several important implications. It should impact the field of corpus-based bilingual lexicography, where a practically usable encoding bilingual dictionary has been long awaited but yet to be seen. The study can also be readily extendable to other fields such as translation studies, second language acquisition and cognitive linguistics – particularly in the area of construction grammar.
| Date of Award | 16 Feb 2015 |
|---|
| Original language | English |
|---|
| Awarding Institution | - City University of Hong Kong
|
|---|
| Supervisor | Chengyu Alex FANG (Supervisor) |
|---|
- English language
- Japanese language
- Computational linguistics
- Data processing
A corpus-based study of encoding processes from the Japanese language to the English language
HIRATA, M. (Author). 16 Feb 2015
Student thesis: Doctoral Thesis