Connecting Phrase Based Statistical Machine Translation Adaptation.

Rui Wang,Hai Zhao,Bao-Liang Lu,Masao Utiyama,Eiichro Sumita
DOI: https://doi.org/10.48550/arxiv.1607.08693
2016-01-01
Abstract:Although more additional corpora are now available for Statistical Machine Translation (SMT), only the ones which belong to the same or similar domains of the original corpus can indeed enhance SMT performance directly. A series of SMT adaptation methods have been proposed to select these similar-domain data, and most of them focus on sentence selection. In comparison, phrase is a smaller and more fine grained unit for data selection, therefore we propose a straightforward and efficient connecting phrase based adaptation method, which is applied to both bilingual phrase pair and monolingual n-gram adaptation. The proposed method is evaluated on IWSLT/NIST data sets, and the results show that phrase based SMT performances are significantly improved (up to +1.6 in comparison with phrase based SMT baseline system and +0.9 in comparison with existing methods).
What problem does this paper attempt to address?