Grapheme Segmentation Based Mongolian Handwriting Recognition

Daoerji FAN,Guanglai GAO,Huijuan WU
DOI: https://doi.org/10.3969/j.issn.1003-0077.2017.05.010
2017-01-01
Abstract:Hidden Markov Models(HMM ) has strong modeling capabilities for sequence data,and it is widely used in speech recognition and handwriting recognition task.HMM-based Mongolian handwriting recognizers require the data to be analyzed sequentially.According to Mongolian word formation and writing style,it is evident that a Mon-golian word consists of grapheme seamless connected from top to down.The selection of grapheme and segmentation word to grapheme is a preliminary work for handwriting recognition with substantial effects on recognition accuracy. In this paper,according to knowledge of syllables and coding,we collect a Mongolian letters set of 1171 letters. The long grapheme set which contain 378 grapheme is then extracted from letters set by correlation process and HMM based sorting method.The short grapheme set which contain 50 shapes is extracted from long grapheme set via decompose long grapheme by hands.We present an algorithm to decompose a word to grapheme by two layers mapping.Experimental results show that the short grapheme get better performance than long grapheme.
What problem does this paper attempt to address?