Abstract
In this study, we propose to use two corpus-driven linguistic approaches for a sense prediction study. We will concentrate on the character similarity clustering approach and concept similarity clustering approach to predict the senses of non-assigned words by using corpora and tools, such as Chinese Gigaword Corpus, and HowNet. In this study, we would then like to evaluate their predictions via the sense divisions of Chinese Wordnet (CWN) and Xiandai Hanyu Cidian (Xian Han). Using these corpora, we will determine their clusters of our four target words ---- chi1 “eat”, wan2 “play”, huan4 “change” and shao1 “burn” in order to predict their all possible senses and evaluate them. This requirement will demonstrate the visibility of the corpus-based approaches.
Original language | English |
---|---|
Title of host publication | Proceedings of the 11th Chinese Lexical Semantic Workshop, CLSW 2010 |
Publisher | Soochow University |
Pages | 239-246 |
Number of pages | 8 |
Publication status | Published - May 2010 |
Event | The 11th Chinese Lexical Semantic Workshop, CLSW 2010 - Suzhou, China Duration: 21 May 2010 → 23 May 2010 |
Conference
Conference | The 11th Chinese Lexical Semantic Workshop, CLSW 2010 |
---|---|
Country/Territory | China |
Period | 21/05/10 → 23/05/10 |
User-Defined Keywords
- Lexical ambiguity
- Sense prediction
- Corpus-based approach
- Character similarity clustering approach
- Concept similarity clustering approach
- Evaluation