Abstract:Semitic morphologically-rich languages (MRLs) are characterized by extreme word ambiguity. Because most vowels are omitted in standard texts, many of the words are homographs with multiple possible analyses, each with a different pronunciation and different morphosyntactic properties. This ambiguity goes beyond word-sense disambiguation (WSD), and may include token segmentation into multiple word units. Previous research on MRLs claimed that standardly trained pre-trained language models (PLMs) based on word-pieces may not sufficiently capture the internal structure of such tokens in order to distinguish between these analyses. Taking Hebrew as a case study, we investigate the extent to which Hebrew homographs can be disambiguated and analyzed using PLMs. We evaluate all existing models for contextualized Hebrew embeddings on a novel Hebrew homograph challenge sets that we deliver. Our empirical results demonstrate that contemporary Hebrew contextualized embeddings outperform non-contextualized embeddings; and that they are most effective for disambiguating segmentation and morphosyntactic features, less so regarding pure word-sense disambiguation. We show that these embeddings are more effective when the number of word-piece splits is limited, and they are more effective for 2-way and 3-way ambiguities than for 4-way ambiguity. We show that the embeddings are equally effective for homographs of both balanced and skewed distributions, whether calculated as masked or unmasked tokens. Finally, we show that these embeddings are as effective for homograph disambiguation with extensive supervised training as with a few-shot setup.

Analyzing the Attention Heads for Pronoun Disambiguation in Context-aware Machine Translation Models

Do Context-Aware Translation Models Pay the Right Attention?

Mention Attention for Pronoun Translation

An Analysis of Attention Mechanisms: The Case of Word Sense Disambiguation in Neural Machine Translation

Comparison of the Effects of Attention Mechanism on Translation Tasks of Different Lengths of Ambiguous Words

Interrogating the Explanatory Power of Attention in Neural Machine Translation

A Large-Scale Test Set for the Evaluation of Context-Aware Pronoun Translation in Neural Machine Translation

Evaluating Pronominal Anaphora in Machine Translation: An Evaluation Measure and a Test Suite

Position-aware Attention for Enhancing the Machine Comprehension

Learning When to Attend for Neural Machine Translation

On The Alignment Problem In Multi-Head Attention-Based Neural Machine Translation

A sort Approach for Anaphora Resolution of Chinese Personal Pronoun Based on Machine Learning Method

Context-Aware Neural Machine Translation Learns Anaphora Resolution

Do Pretrained Contextual Language Models Distinguish between Hebrew Homograph Analyses?

A Closer Look at Transformer Attention for Multilingual Translation.

Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned

Handling Homographs in Neural Machine Translation

Zero Pronoun Resolution with Attention-based Neural Network.

Attention Mechanism and Context Modeling System for Text Mining Machine Translation

Does Context Help Mitigate Gender Bias in Neural Machine Translation?

Pay Better Attention to Attention: Head Selection in Multilingual and Multi-Domain Sequence Modeling