Abstract:The advent of Multilingual Language Models (MLLMs) and Large Language Models has spawned innovation in many areas of natural language processing. Despite the exciting potential of this technology, its impact on developing high-quality Machine Translation (MT) outputs for low-resource languages remains relatively under-explored. Furthermore, an open-source application, dedicated to both fine-tuning MLLMs and managing the complete MT workflow for low-resources languages, remains unavailable. We aim to address these imbalances through the development of adaptMLLM, which streamlines all processes involved in the fine-tuning of MLLMs for MT. This open-source application is tailored for developers, translators, and users who are engaged in MT. An intuitive interface allows for easy customisation of hyperparameters, and the application offers a range of metrics for model evaluation and the capability to deploy models as a translation service directly within the application. As a multilingual tool, we used adaptMLLM to fine-tune models for two low-resource language pairs: English to Irish (EN$\leftrightarrow$GA) and English to Marathi (EN$\leftrightarrow$MR). Compared with baselines from the LoResMT2021 Shared Task, the adaptMLLM system demonstrated significant improvements. In the EN$\rightarrow$GA direction, an improvement of 5.2 BLEU points was observed and an increase of 40.5 BLEU points was recorded in the GA$\rightarrow$EN direction. Significant improvements in the translation performance of the EN$\leftrightarrow$MR pair were also observed notably in the MR$\rightarrow$EN direction with an increase of 21.3 BLEU points. Finally, a fine-grained human evaluation of the MLLM output on the EN$\rightarrow$GA pair was conducted using the Multidimensional Quality Metrics and Scalar Quality Metrics error taxonomies. The application and models are freely available.

LlamaTurk: Adapting Open-Source Generative Large Language Models for Low-Resource Language

Bridging the Bosphorus: Advancing Turkish Large Language Models through Strategies for Low-Resource Language Adaptation and Benchmarking

Efficient Adaptation: Enhancing Multilingual Models for Low-Resource Language Translation

Efficiently Adapting Pretrained Language Models To New Languages

Targeted Multilingual Adaptation for Low-resource Language Families

Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training

Better to Ask in English: Evaluation of Large Language Models on English, Low-resource and Cross-Lingual Settings

Challenges in Adapting Multilingual LLMs to Low-Resource Languages using LoRA PEFT Tuning

adaptMLLM: Fine-Tuning Multilingual Language Models on Low-Resource Languages with Integrated LLM Playgrounds

Bilingual Adaptation of Monolingual Foundation Models

Improving Language Models Trained on Translated Data with Continual Pre-Training and Dictionary Learning Analysis

Open Generative Large Language Models for Galician

A Three-Pronged Approach to Cross-Lingual Adaptation with Multilingual LLMs

Optimizing Low-Resource Language Model Training: Comprehensive Analysis of Multi-Epoch, Multi-Lingual, and Two-Stage Approaches

Linguistically-Informed Multilingual Instruction Tuning: Is There an Optimal Set of Languages to Tune?

Performance of Recent Large Language Models for a Low-Resourced Language

Quality or Quantity? On Data Scale and Diversity in Adapting Large Language Models for Low-Resource Translation

Exploring Design Choices for Building Language-Specific LLMs

Do Large Language Models Speak All Languages Equally? A Comparative Study in Low-Resource Settings

SambaLingo: Teaching Large Language Models New Languages

Generative Model for Less-Resourced Language with 1 billion parameters