Automatic construction of English/Chinese parallel corpora

Christopher C. Yang,Kar Wing Li

DOI: https://doi.org/10.1002/asi.10261

2003-01-01

Journal of the American Society for Information Science and Technology

Abstract:As the demand for global information increases significantly, multilingual corpora has become a valuable linguistic resource for applications to cross‐lingual information retrieval and natural language processing. In order to cross the boundaries that exist between different languages, dictionaries are the most typical tools. However, the general‐purpose dictionary is less sensitive in both genre and domain. It is also impractical to manually construct tailored bilingual dictionaries or sophisticated multilingual thesauri for large applications. Corpus‐based approaches, which do not have the limitation of dictionaries, provide a statistical translation model with which to cross the language boundary. There are many domain‐specific parallel or comparable corpora that are employed in machine translation and cross‐lingual information retrieval. Most of these are corpora between Indo‐European languages, such as English/French and English/Spanish. The Asian/Indo‐European corpus, especially English/Chinese corpus, is relatively sparse. The objective of the present research is to construct English/Chinese parallel corpus automatically from the World Wide Web. In this paper, an alignment method is presented which is based on dynamic programming to identify the one‐to‐one Chinese and English title pairs. The method includes alignment at title level, word level and character level. The longest common subsequence (LCS) is applied to find the most reliable Chinese translation of an English word. As one word for a language may translate into two or more words repetitively in another language, the edit operation, deletion, is used to resolve redundancy. A score function is then proposed to determine the optimal title pairs. Experiments have been conducted to investigate the performance of the proposed method using the daily press release articles by the Hong Kong SAR government as the test bed. The precision of the result is 0.998 while the recall is 0.806. The release articles and speech articles, published by Hongkong & Shanghai Banking Corporation Limited, are also used to test our method, the precision is 1.00, and the recall is 0.948.

What problem does this paper attempt to address?

The problem that this paper attempts to solve is to construct an English - Chinese parallel corpus to overcome the barriers between different languages in the context of a significant increase in global information needs. Specifically, the goal of the paper is to automatically construct an English - Chinese parallel corpus from the World Wide Web. This is mainly because: 1. **Growth in the demand for multilingual information**: With the increase in Internet users, especially the rapid growth of non - English - speaking users, the demand for cross - language information retrieval has grown exponentially. 2. **Insufficiency of existing resources**: Existing bilingual dictionaries are not sensitive enough in terms of genre and domain, and it is unrealistic to manually construct customized bilingual dictionaries or complex multilingual dictionaries for large - scale applications. 3. **Special challenges between Asian and Indo - European languages**: Compared with European languages, the machine translation technology maturity between Asian languages (such as Chinese) and Indo - European languages is lower, parallel or comparable corpora are relatively scarce, and there are significant grammatical differences. To address these challenges, the paper proposes a dynamic - programming - based method to identify one - to - one matches between Chinese and English titles. This method includes the following steps: - **Title - level alignment**: Find the most reliable Chinese translation through the Longest Common Subsequence (LCS) algorithm. - **Word - level and character - level alignment**: Considering that a word in one language may be translated into multiple words in another language, use the deletion operation to solve the redundancy problem. - **Scoring function**: Propose a scoring function to determine the best title pairs. The experimental results show that the precision of this method on government press releases is 0.998 and the recall rate is 0.806; on HSBC's press releases and speeches, the precision is 1.00 and the recall rate is 0.948. These results verify the effectiveness of this method.

Automatic construction of English/Chinese parallel corpora

Development of Translation Database based on Chinese-English parallel corpora

Chinese-English Parallel Corpus Construction And Its Application

Constructing of a large-scale Chinese-English parallel corpus

Design of New Word Retrieval Algorithm for Chinese-English Bilingual Parallel Corpus

UM-Corpus: A Large English-Chinese Parallel Corpus for Statistical Machine Translation

Aligning a Parallel English-Chinese Corpus Statistically with Lexical Criteria

Extraction of translation unit from Chinese-English parallel corpora

The Construction of a Chinese-English Patent Parallel Corpus

Automatic Alignment of English-Chinese Bilingual Texts of CNS News

Research of English-Chinese Alignment at Word Granularity on Parallel Corpora

A Study of Business English Translation Skills Based on Parallel Corpus

NEJM-enzh: A Parallel Corpus for English-Chinese Translation in the Biomedical Domain

Chinese-Portuguese Machine Translation: A Study on Building Parallel Corpora from Comparable Texts

A Japanese-Chinese Parallel Corpus Using Crowdsourcing for Web Mining

Alignment and Extraction of Bilingual Legal Terminology from Context Profiles

Automatic Translating Between Ancient Chinese and Contemporary Chinese with Limited Aligned Corpora.

Building a Large English-Chinese Parallel Corpus from Comparable Patents and Its Experimental Application to SMT

Bilingual Terminology Extraction from Comparable E-Commerce Corpora

Building a Parallel Corpus for English Translation Teaching Based on Computer-Aided Translation Software

WCC-JC 2.0: A Web-Crawled and Manually Aligned Parallel Corpus for Japanese-Chinese Neural Machine Translation