Multilingual Contextualization of Large Language Models for Document-Level Machine Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ramos, Miguel Moura, Fernandes, Patrick, Agrawal, Sweta, Martins, André F. T.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916922563493888
author Ramos, Miguel Moura
Fernandes, Patrick
Agrawal, Sweta
Martins, André F. T.
author_facet Ramos, Miguel Moura
Fernandes, Patrick
Agrawal, Sweta
Martins, André F. T.
contents Large language models (LLMs) have demonstrated strong performance in sentence-level machine translation, but scaling to document-level translation remains challenging, particularly in modeling long-range dependencies and discourse phenomena across sentences and paragraphs. In this work, we propose a method to improve LLM-based long-document translation through targeted fine-tuning on high-quality document-level data, which we curate and introduce as DocBlocks. Our approach supports multiple translation paradigms, including direct document-to-document and chunk-level translation, by integrating instructions both with and without surrounding context. This enables models to better capture cross-sentence dependencies while maintaining strong sentence-level translation performance. Experimental results show that incorporating multiple translation paradigms improves document-level translation quality and inference speed compared to prompting and agent-based methods.
format Preprint
id arxiv_https___arxiv_org_abs_2504_12140
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multilingual Contextualization of Large Language Models for Document-Level Machine Translation
Ramos, Miguel Moura
Fernandes, Patrick
Agrawal, Sweta
Martins, André F. T.
Computation and Language
Large language models (LLMs) have demonstrated strong performance in sentence-level machine translation, but scaling to document-level translation remains challenging, particularly in modeling long-range dependencies and discourse phenomena across sentences and paragraphs. In this work, we propose a method to improve LLM-based long-document translation through targeted fine-tuning on high-quality document-level data, which we curate and introduce as DocBlocks. Our approach supports multiple translation paradigms, including direct document-to-document and chunk-level translation, by integrating instructions both with and without surrounding context. This enables models to better capture cross-sentence dependencies while maintaining strong sentence-level translation performance. Experimental results show that incorporating multiple translation paradigms improves document-level translation quality and inference speed compared to prompting and agent-based methods.
title Multilingual Contextualization of Large Language Models for Document-Level Machine Translation
topic Computation and Language
url https://arxiv.org/abs/2504.12140