Transformed Protoform Reconstruction

Young Min Kim,Kalvin Chang,Chenxuan Cui,David Mortensen
2023-07-06
Abstract:Protoform reconstruction is the task of inferring what morphemes or words appeared like in the ancestral languages of a set of daughter languages. Meloni et al. (2021) achieved the state-of-the-art on Latin protoform reconstruction with an RNN-based encoder-decoder with attention model. We update their model with the state-of-the-art seq2seq model: the Transformer. Our model outperforms their model on a suite of different metrics on two different datasets: their Romance data of 8,000 cognates spanning 5 languages and a Chinese dataset (Hou 2004) of 800+ cognates spanning 39 varieties. We also probe our model for potential phylogenetic signal contained in the model. Our code is publicly available at <a class="link-external link-https" href="https://github.com/cmu-llab/acl-2023" rel="external noopener nofollow">this https URL</a>.
Computation and Language
What problem does this paper attempt to address?