Abstract:In this article, we conduct an empirical investigation of translation divergences between Chinese and English relying on a parallel treebank. To do this, we first devise a hierarchical alignment scheme where Chinese and English parse trees are aligned in a way that eliminates conflicts and redundancies between word alignments and syntactic parses to prevent the generation of spurious translation divergences. Using this Hierarchically Aligned Chinese–English Parallel Treebank (HACEPT), we are able to semi-automatically identify and categorize the translation divergences between the two languages and quantify each type of translation divergence. Our results show that the translation divergences are much broader than described in previous studies that are largely based on anecdotal evidence and linguistic knowledge. The distribution of the translation divergences also shows that some high-profile translation divergences that motivate previous research are actually very rare in our data, whereas other translation divergences that have previously received little attention actually exist in large quantities. We also show that HACEPT allows the extraction of syntax-based translation rules, most of which are expressive enough to capture the translation divergences, and point out that the syntactic annotation in existing treebanks is not optimal for extracting such translation rules. We also discuss the implications of our study for attempts to bridge translation divergences by devising shared semantic representations across languages. Our quantitative results lend further support to the observation that although it is possible to bridge some translation divergences with semantic representations, other translation divergences are open-ended, thus building a semantic representation that captures all possible translation divergences may be impractical.

That was the last straw, we need more: Are Translation Systems Sensitive to Disambiguating Context?

When a Good Translation is Wrong in Context: Context-Aware Machine Translation Improves on Deixis, Ellipsis, and Lexical Cohesion

When Does Translation Require Context? A Data-driven, Multilingual Exploration

Towards Effective Disambiguation for Machine Translation with Large Language Models

Using Language Models to Disambiguate Lexical Choices in Translation

Shades of meaning: Uncovering the geometry of ambiguous word representations through contextualised language models

Processing of Translation-Ambiguous Words by Chinese–English Bilinguals in Sentence Context

Speakers Fill Lexical Semantic Gaps with Context

How sensitive are translation systems to extra contexts? Mitigating gender bias in Neural Machine Translation models through relevant contexts

Towards Accurate Translation via Semantically Appropriate Application of Lexical Constraints

Improving Word Sense Disambiguation in Neural Machine Translation with Salient Document Context

Escaping the sentence-level paradigm in machine translation

Translate Meanings, Not Just Words: IdiomKB's Role in Optimizing Idiomatic Translation with Language Models

Can LLMs assist with Ambiguity? A Quantitative Evaluation of various Large Language Models on Word Sense Disambiguation

Analyzing Context Utilization of LLMs in Document-Level Translation

We're Afraid Language Models Aren't Modeling Ambiguity

Evaluating Contextualized Representations of (Spanish) Ambiguous Words: A New Lexical Resource and Empirical Analysis

Analyzing Context Contributions in LLM-based Machine Translation

Identifying Context-Dependent Translations for Evaluation Set Production

Paying More Attention to Source Context: Mitigating Unfaithful Translations from Large Language Model

Translation Divergences in Chinese–English Machine Translation: an Empirical Investigation