Abstract:The quality of neural machine translation can be improved by leveraging additional monolingual resources to create synthetic training data. Source-side monolingual data can be (forward-)translated into the target language for self-training; target-side monolingual data can be back-translated. It has been widely reported that back-translation delivers superior results, but could this be due to artefacts in the test sets? We perform a case study using French-English news translation task and separate test sets based on their original languages. We show that forward translation delivers superior gains in terms of BLEU on sentences that were originally in the source language, complementing previous studies which show large improvements with back-translation on sentences that were originally in the target language. To better understand when and why forward and back-translation are effective, we study the role of domains, translationese, and noise. While translationese effects are well known to influence MT evaluation, we also find evidence that news data from different languages shows subtle domain differences, which is another explanation for varying performance on different portions of the test set. We perform additional low-resource experiments which demonstrate that forward translation is more sensitive to the quality of the initial translation system than back-translation, and tends to perform worse in low-resource settings.

Training on Synthetic Noise Improves Robustness to Natural Noise in Machine Translation

Synthetic and Natural Noise Both Break Neural Machine Translation

Improving Robustness of Machine Translation with Synthetic Noise

Robust Neural Machine Translation: Modeling Orthographic and Interpunctual Variation

Data Noising as Smoothing in Neural Network Language Models

Domain, Translationese and Noise in Synthetic Data for Neural Machine Translation

Did Translation Models Get More Robust Without Anyone Even Noticing?

Addressing the Vulnerability of NMT in Input Perturbations

Improving the Robustness of Speech Translation

Towards Robust Transcription: Exploring Noise Injection Strategies for Training Data Augmentation

Robust Neural Machine Translation for Clean and Noisy Speech Transcripts

Alternated Training with Synthetic and Authentic Data for Neural Machine Translation

Towards Robust Neural Machine Translation

Exploiting Monolingual Data at Scale for Neural Machine Translation.

Robustness Enhancement in Neural Networks with Alpha-Stable Training Noise

On Synthetic Data for Back Translation

Noisy UGC Translation at the Character Level: Revisiting Open-Vocabulary Capabilities and Robustness of Char-Based Models

Enhancing Model Robustness Via Lexical Distilling

Error Norm Truncation: Robust Training in the Presence of Data Noise for Text Generation Models

Non-Fluent Synthetic Target-Language Data Improve Neural Machine Translation

NAT: Noise-Aware Training for Robust Neural Sequence Labeling