Abstract:Word reordering is one of the most difficult aspects of statistical machine translation (SMT), and an important factor of its quality and efficiency. Despite the vast amount of research published to date, the interest of the community in this problem has not decreased, and no single method appears to be strongly dominant across language pairs. Instead, the choice of the optimal approach for a new translation task still seems to be mostly driven by empirical trials. To orientate the reader in this vast and complex research area, we present a comprehensive survey of word reordering viewed as a statistical modeling challenge and as a natural language phenomenon. The survey describes in detail how word reordering is modeled within different string-based and tree-based SMT frameworks and as a stand-alone task, including systematic overviews of the literature in advanced reordering modeling. We then question why some approaches are more successful than others in different language pairs. We argue that, besides measuring the amount of reordering, it is important to understand which kinds of reordering occur in a given language pair. To this end, we conduct a qualitative analysis of word reordering phenomena in a diverse sample of language pairs, based on a large collection of linguistic knowledge. Empirical results in the SMT literature are shown to support the hypothesis that a few linguistic facts can be very useful to anticipate the reordering characteristics of a language pair and to select the SMT framework that best suits them.

Learning local word reorderings for hierarchical phrase-based statistical machine translation

Learning Word Reorderings For Hierarchical Phrase-Based Statistical Machine Translation

A Ranking-based Approach to Word Reordering for Statistical Machine Translation.

A Lexicalized Reordering Model for Hierarchical Phrase-based Translation

Reordering with Source Language Collocations.

3 Reordering Model with Source Language Collocations

A Novel Word Reordering Method For Statistical Machine Translation

Graph-based Lexicalized Reordering Models for Statistical Machine Translation

Learning Lexicalized Reordering Models from Reordering Graphs.

A Neural Reordering Model for Phrase-based Translation.

Lexicalized Reordering In Multiple-Graph Based Statistical Machine Translation

Lexical Reordering for Hierarchical Phrase-based Translation

Word-Level Reordering Model for Phrase-Based SMT

Neural Machine Translation with Reordering Embeddings.

Phrase Based Language Model for Statistical Machine Translation: Empirical Study

A Survey of Word Reordering in Statistical Machine Translation: Computational Models and Language Phenomena

A novel dependency based word-level reordering model for phrased-based translation

Head-modifier relation based non-lexical reordering model for phrase-based translation

An Improved Hierarchical Phrase Based Machine Translation Model

A Probabilistic Approach to Syntax-based Reordering for Statistical Machine Translation.

To Swap or Not to Swap? Exploiting Dependency Word Pairs for Reordering in Statistical Machine Translation.