Multilingual Contextualization of Large Language Models for Document-Level Machine Translation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916922563493888 |
|---|---|
| author | Ramos, Miguel Moura Fernandes, Patrick Agrawal, Sweta Martins, André F. T. |
| author_facet | Ramos, Miguel Moura Fernandes, Patrick Agrawal, Sweta Martins, André F. T. |
| contents | Large language models (LLMs) have demonstrated strong performance in sentence-level machine translation, but scaling to document-level translation remains challenging, particularly in modeling long-range dependencies and discourse phenomena across sentences and paragraphs. In this work, we propose a method to improve LLM-based long-document translation through targeted fine-tuning on high-quality document-level data, which we curate and introduce as DocBlocks. Our approach supports multiple translation paradigms, including direct document-to-document and chunk-level translation, by integrating instructions both with and without surrounding context. This enables models to better capture cross-sentence dependencies while maintaining strong sentence-level translation performance. Experimental results show that incorporating multiple translation paradigms improves document-level translation quality and inference speed compared to prompting and agent-based methods. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_12140 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Multilingual Contextualization of Large Language Models for Document-Level Machine Translation Ramos, Miguel Moura Fernandes, Patrick Agrawal, Sweta Martins, André F. T. Computation and Language Large language models (LLMs) have demonstrated strong performance in sentence-level machine translation, but scaling to document-level translation remains challenging, particularly in modeling long-range dependencies and discourse phenomena across sentences and paragraphs. In this work, we propose a method to improve LLM-based long-document translation through targeted fine-tuning on high-quality document-level data, which we curate and introduce as DocBlocks. Our approach supports multiple translation paradigms, including direct document-to-document and chunk-level translation, by integrating instructions both with and without surrounding context. This enables models to better capture cross-sentence dependencies while maintaining strong sentence-level translation performance. Experimental results show that incorporating multiple translation paradigms improves document-level translation quality and inference speed compared to prompting and agent-based methods. |
| title | Multilingual Contextualization of Large Language Models for Document-Level Machine Translation |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2504.12140 |