Inducing Bilingual Lexica from Non-Parallel Data with Earth Mover's Distance Regularization.

Meng Zhang,Yang Liu,Huanbo Luan,Yiqun Liu,Maosong Sun
2016-01-01
Abstract:Being able to induce word translations from non-parallel data is often a prerequisite for cross-lingual processing in resource-scarce languages and domains. Previous endeavors typically simplify this task by imposing the one-to-one translation assumption, which is too strong to hold for natural languages. We remove this constraint by introducing the Earth Mover’s Distance into the training of bilingual word embeddings. In this way, we take advantage of its capability to handle multiple alternative word translations in a natural form of regularization. Our approach shows significant and consistent improvements across four language pairs. We also demonstrate that our approach is particularly preferable in resource-scarce settings as it only requires a minimal seed lexicon.
What problem does this paper attempt to address?