Abstract:In-image machine translation (IIMT) aims to translate an image containing texts in source language into an image containing translations in target language. In this regard, conventional cascaded methods suffer from issues such as error propagation, massive parameters, and difficulties in deployment and retaining visual characteristics of the input image. Thus, constructing end-to-end models has become an option, which, however, faces two main challenges: 1) the huge modeling burden, as it is required to simultaneously learn alignment across languages and preserve the visual characteristics of the input image; 2) the difficulties of directly predicting excessively lengthy pixel sequences. In this paper, we propose \textit{Translatotron-V(ision)}, an end-to-end IIMT model consisting of four modules. In addition to an image encoder, and an image decoder, our model contains a target text decoder and an image tokenizer. Among them, the target text decoder is used to alleviate the language alignment burden, and the image tokenizer converts long sequences of pixels into shorter sequences of visual tokens, preventing the model from focusing on low-level visual features. Besides, we present a two-stage training framework for our model to assist the model in learning alignment across modalities and languages. Finally, we propose a location-aware evaluation metric called Structure-BLEU to assess the translation quality of the generated images. Experimental results demonstrate that our model achieves competitive performance compared to cascaded models with only 70.9\% of parameters, and significantly outperforms the pixel-level end-to-end IIMT model.

Improving End-to-End Text Image Translation From the Auxiliary Text Translation Task

Cross-Lingual Text Image Recognition Via Multi-Task Sequence to Sequence Learning.

Improving Speech Translation by Understanding and Learning from the Auxiliary Text Translation Task

Adaptive multi-task learning for speech to text translation

Rethinking and Improving Multi-task Learning for End-to-end Speech Translation

Exploring Better Text Image Translation with Multimodal Codebook

Modal Contrastive Learning based End-to-End Text Image Machine Translation

Translation-Enhanced Multilingual Text-to-Image Generation

HybridVocab: Towards Multi-Modal Machine Translation Via Multi-Aspect Alignment

AnyTrans: Translate AnyText in the Image with Large Scale Models

MIT-10M: A Large Scale Parallel Corpus of Multilingual Image Translation

Transductive Auxiliary Task Self-Training for Neural Multi-Task Models

Imagination improves Multimodal Translation

Cross-Modal Multi-Tasking for Speech-to-Text Translation via Hard Parameter Sharing

Translatotron-V(ison): An End-to-End Model for In-Image Machine Translation

Improving General Text Embedding Model: Tackling Task Conflict and Data Imbalance through Model Merging

A multitask co-training framework for improving speech translation by leveraging speech recognition and machine translation tasks

A General Multi-Task Learning Framework to Leverage Text Data for Speech to Text Tasks

Multilingual Multimodal Learning with Machine Translated Text

Beyond Filtering: Adaptive Image-Text Quality Enhancement for MLLM Pretraining

Multi-Task Learning for Front-End Text Processing in TTS