Abstract:As the demand for global information increases significantly, multilingual corpora has become a valuable linguistic resource for applications to cross‐lingual information retrieval and natural language processing. In order to cross the boundaries that exist between different languages, dictionaries are the most typical tools. However, the general‐purpose dictionary is less sensitive in both genre and domain. It is also impractical to manually construct tailored bilingual dictionaries or sophisticated multilingual thesauri for large applications. Corpus‐based approaches, which do not have the limitation of dictionaries, provide a statistical translation model with which to cross the language boundary. There are many domain‐specific parallel or comparable corpora that are employed in machine translation and cross‐lingual information retrieval. Most of these are corpora between Indo‐European languages, such as English/French and English/Spanish. The Asian/Indo‐European corpus, especially English/Chinese corpus, is relatively sparse. The objective of the present research is to construct English/Chinese parallel corpus automatically from the World Wide Web. In this paper, an alignment method is presented which is based on dynamic programming to identify the one‐to‐one Chinese and English title pairs. The method includes alignment at title level, word level and character level. The longest common subsequence (LCS) is applied to find the most reliable Chinese translation of an English word. As one word for a language may translate into two or more words repetitively in another language, the edit operation, deletion, is used to resolve redundancy. A score function is then proposed to determine the optimal title pairs. Experiments have been conducted to investigate the performance of the proposed method using the daily press release articles by the Hong Kong SAR government as the test bed. The precision of the result is 0.998 while the recall is 0.806. The release articles and speech articles, published by Hongkong & Shanghai Banking Corporation Limited, are also used to test our method, the precision is 1.00, and the recall is 0.948.

Construct trilingual parallel corpus on demand

Automatic construction of English/Chinese parallel corpora

Chinese-English Parallel Corpus Construction And Its Application

UM-Corpus: A Large English-Chinese Parallel Corpus for Statistical Machine Translation

Constructing of a large-scale Chinese-English parallel corpus

Automatic English-Chinese parallel corpus acquisition and sentences extraction

Development of Translation Database based on Chinese-English parallel corpora

The Cultivation of a Chinese-English-Japanese Trilingual Parallel Corpus from Comparable Patents

Construction and Processing of a Parallel Corpus for Tang Poetry and Song Lyrics

Inflating a Small Parallel Corpus into a Large Quasi-parallel Corpus Using Monolingual Data for Chinese-Japanese Machine Translation.

Automatic Construction of Discourse Corpora for Dialogue Translation

Recent Developments in Chinese Corpus Research

Chinese-Portuguese Machine Translation: A Study on Building Parallel Corpora from Comparable Texts

Bilingual Corpus Mining and Multistage Fine-Tuning for Improving Machine Translation of Lecture Transcripts

Korean-Centered Cross-Lingual Parallel Sentence Corpus Construction Experiment

A Bilingual Corpus in the Legal Domain and its Applications

A Japanese-Chinese Parallel Corpus Using Crowdsourcing for Web Mining

Building a Large English-Chinese Parallel Corpus from Comparable Patents and Its Experimental Application to SMT

The translation teaching platform based on multilingual corpora of Xi Jinping: The Governance of China: Design, resources and applications

NEJM-enzh: A Parallel Corpus for English-Chinese Translation in the Biomedical Domain

A Construction Method of Multilingual Comparable Corpus in the Background of Artificial Intelligence and Internet of Things