Pronunciation-Enhanced Chinese Word Embedding

Qinjuan Yang,Haoran Xie,Gary Cheng,Fu Lee Wang,Yanghui Rao
DOI: https://doi.org/10.1007/s12559-021-09850-9
IF: 4.89
2021-02-22
Cognitive Computation
Abstract:Abstract Chinese word embeddings have recently garnered considerable attention. Chinese characters and their sub-character components, which contain rich semantic information, are incorporated to learn Chinese word embeddings. Chinese characters can represent a combination of meaning, structure, and pronunciation. However, existing embedding learning methods focus on the structure and meaning of Chinese characters. In this study, we aim to develop an embedding learning method that can make complete use of the information represented by Chinese characters, including phonology, morphology, and semantics. Specifically, we propose a pronunciation-enhanced Chinese word embedding learning method, where the pronunciations of context characters and target characters are simultaneously encoded into the embeddings. Evaluation of word similarity, word analogy reasoning, text classification, and sentiment analysis validate the effectiveness of our proposed method.
computer science, artificial intelligence,neurosciences
What problem does this paper attempt to address?
The paper aims to address the issue of capturing character polysemy in Chinese word embedding methods and proposes an improved method that combines phonological, morphological, and semantic features. Specifically, the paper introduces a method called Pronunciation-Enhanced Chinese Word Embedding (PCWE). This method predicts target words by integrating the pronunciation information of Chinese characters and their subcomponents, thereby improving the quality of word embeddings. Experimental results show that this method achieves better performance in tasks such as word similarity, word analogy reasoning, text classification, and sentiment analysis, especially demonstrating stronger capabilities in handling polyphonic characters.