Improved statistical machine translation using monolingual paraphrases

Preslav Nakov
DOI: https://doi.org/10.48550/arXiv.2109.15119
2021-09-26
Abstract:We propose a novel monolingual sentence paraphrasing method for augmenting the training data for statistical machine translation systems "for free" -- by creating it from data that is already available rather than having to create more aligned data. Starting with a syntactic tree, we recursively generate new sentence variants where noun compounds are paraphrased using suitable prepositions, and vice-versa -- preposition-containing noun phrases are turned into noun compounds. The evaluation shows an improvement equivalent to 33%-50% of that of doubling the amount of training data.
Computation and Language,Artificial Intelligence
What problem does this paper attempt to address?