Corpus-Based Extraction of Collocations in Chinese

Hui Wang,Dong-Hong Ji
DOI: https://doi.org/10.1109/wiiat.2008.72
2008-01-01
Abstract:Collocation, i.e. the sequences of certain words which habitually co-occur, plays an essential part in human language. The present study is intending to identify the detailed classification and typical features of collocations in Chinese language, and explore a new computer-assistant way for extraction and representation of Chinese collocations. The investigation is based on the largest and only Singapore Chinese corpus (SCC), of which 20 million words have been analysed. The central novel idea of this research is the combination of dictionary, language rules and statistic data in automatic collocation extraction. So far, this method has not been proposed.
What problem does this paper attempt to address?