Constructing of a large-scale Chinese-English parallel corpus

Le Sun,Song Xue,Weimin Qu,Xiaofeng Wang,Yufang Sun
DOI: https://doi.org/10.3115/1118759.1118770
2002-01-01
Abstract:This paper describes the constructing of a large-scale (above 500,000 pair sentences) Chinese-English parallel corpus. The current status of Chinese corpora is overviewed with the emphasis on parallel corpus. The XML coding principles for Chinese--English parallel corpus are discussed. The sentence alignment algorithm used in this project is described with a computer-aided checking processing. Finally, we show the design of the concordance of the parallel corpus and the prospect to further development.
What problem does this paper attempt to address?