Abstract:The large language model (LLM) has garnered significant attention due to its in-context learning mechanisms and emergent capabilities. The research community has conducted several pilot studies to apply LLMs to machine translation tasks and evaluate their performance from diverse perspectives. However, previous research has primarily focused on the LLM itself and has not explored human intervention in the inference process of LLM. The characteristics of LLM, such as in-context learning and prompt engineering, closely mirror human cognitive abilities in language tasks, offering an intuitive solution for human-in-the-loop generation. In this study, we propose a human-in-the-loop pipeline that guides LLMs to produce customized outputs with revision instructions. The pipeline initiates by prompting the LLM to produce a draft translation, followed by the utilization of automatic retrieval or human feedback as supervision signals to enhance the LLM's translation through in-context learning. The human-machine interactions generated in this pipeline are also stored in an external database to expand the in-context retrieval database, enabling us to leverage human supervision in an offline setting. We evaluate the proposed pipeline using GPT-3.5-turbo API on five domain-specific benchmarks for German-English translation. The results demonstrate the effectiveness of the pipeline in tailoring in-domain translations and improving translation performance compared to direct translation. Additionally, we discuss the results from the following perspectives: 1) the effectiveness of different in-context retrieval methods; 2) the construction of a retrieval database under low-resource scenarios; 3) the observed domains differences; 4) the quantitative analysis of linguistic statistics; and 5) the qualitative analysis of translation cases. The code and data are available at <a class="link-external link-https" href="https://github.com/NLP2CT/HIL-MT/" rel="external noopener nofollow">this https URL</a>.

Ladder: A Model-Agnostic Framework Boosting LLM-based Machine Translation to the Next Level

A Paradigm Shift in Machine Translation: Boosting Translation Performance of Large Language Models

A Novel Paradigm Boosting Translation Capabilities of Large Language Models

ModelGPT: Unleashing LLM's Capabilities for Tailored Model Generation

Multilingual Machine Translation with Large Language Models: Empirical Results and Analysis

Document-Level Machine Translation with Large Language Models

Human-in-the-loop Machine Translation with Large Language Model

LMTuner: An user-friendly and highly-integrable Training Framework for fine-tuning Large Language Models

A Paradigm Shift: The Future of Machine Translation Lies with Large Language Models

Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation

TEaR: Improving LLM-based Machine Translation with Systematic Self-Refinement

Adapting Large Language Models for Document-Level Machine Translation

GPTA: Generative Prompt Tuning Assistant for Synergistic Downstream Neural Network Enhancement with LLMs

CodeLutra: Boosting LLM Code Generation via Preference-Guided Refinement

Exploring Human-Like Translation Strategy with Large Language Models

Enhancing Large Language Model Performance To Answer Questions and Extract Information More Accurately

GPT-4 vs. Human Translators: A Comprehensive Evaluation of Translation Quality Across Languages, Domains, and Expertise Levels

MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models

adaptMLLM: Fine-Tuning Multilingual Language Models on Low-Resource Languages with Integrated LLM Playgrounds

Adaptive Machine Translation with Large Language Models

BigTranslate: Augmenting Large Language Models with Multilingual Translation Capability over 100 Languages