What Does LLM Refinement Actually Improve? A Systematic Study on Document-Level Literary Translation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910216117813248 |
|---|---|
| author | Tan, Shaomu Zhu, Dawei Tran, Ke Denkowski, Michael Trenous, Sony Byrne, Bill Ribeiro, Leonardo Hieber, Felix |
| author_facet | Tan, Shaomu Zhu, Dawei Tran, Ke Denkowski, Michael Trenous, Sony Byrne, Bill Ribeiro, Leonardo Hieber, Felix |
| contents | Iterative self-refinement is a simple inference-time strategy for machine translation: an LLM revises its own translation over multiple inference-time passes. Yet document-scale refinement remains poorly understood: 1) which pipelines work best, 2) what quality dimensions improve, and 3) how refiners behave. In this paper, we present a systematic study of document-level literary translation, covering nine LLMs and seven language pairs. Across nine translation-refinement granularity combinations and five refinement strategies, we find a robust recipe: document-level MT followed by segment-level refinement yields strong and stable improvements. In contrast, document-level refinement often makes fewer edits and leads to smaller or less reliable gains. Beyond granularity, A simple general refinement prompt consistently outperforms error-specific prompting and evaluate-then-refine schemes. Our large-scale human evaluation shows that refinement gains come primarily from fluency, style, and terminology, with limited and less consistent improvements in adequacy. Experiments varying model strength reveal refinement projects outputs toward the refiner's distribution rather than performing targeted error repair. These findings clarify the mechanisms and limitations of current refinement approaches. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_13368 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | What Does LLM Refinement Actually Improve? A Systematic Study on Document-Level Literary Translation Tan, Shaomu Zhu, Dawei Tran, Ke Denkowski, Michael Trenous, Sony Byrne, Bill Ribeiro, Leonardo Hieber, Felix Computation and Language Iterative self-refinement is a simple inference-time strategy for machine translation: an LLM revises its own translation over multiple inference-time passes. Yet document-scale refinement remains poorly understood: 1) which pipelines work best, 2) what quality dimensions improve, and 3) how refiners behave. In this paper, we present a systematic study of document-level literary translation, covering nine LLMs and seven language pairs. Across nine translation-refinement granularity combinations and five refinement strategies, we find a robust recipe: document-level MT followed by segment-level refinement yields strong and stable improvements. In contrast, document-level refinement often makes fewer edits and leads to smaller or less reliable gains. Beyond granularity, A simple general refinement prompt consistently outperforms error-specific prompting and evaluate-then-refine schemes. Our large-scale human evaluation shows that refinement gains come primarily from fluency, style, and terminology, with limited and less consistent improvements in adequacy. Experiments varying model strength reveal refinement projects outputs toward the refiner's distribution rather than performing targeted error repair. These findings clarify the mechanisms and limitations of current refinement approaches. |
| title | What Does LLM Refinement Actually Improve? A Systematic Study on Document-Level Literary Translation |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2605.13368 |