What Does LLM Refinement Actually Improve? A Systematic Study on Document-Level Literary Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tan, Shaomu, Zhu, Dawei, Tran, Ke, Denkowski, Michael, Trenous, Sony, Byrne, Bill, Ribeiro, Leonardo, Hieber, Felix
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910216117813248
author Tan, Shaomu
Zhu, Dawei
Tran, Ke
Denkowski, Michael
Trenous, Sony
Byrne, Bill
Ribeiro, Leonardo
Hieber, Felix
author_facet Tan, Shaomu
Zhu, Dawei
Tran, Ke
Denkowski, Michael
Trenous, Sony
Byrne, Bill
Ribeiro, Leonardo
Hieber, Felix
contents Iterative self-refinement is a simple inference-time strategy for machine translation: an LLM revises its own translation over multiple inference-time passes. Yet document-scale refinement remains poorly understood: 1) which pipelines work best, 2) what quality dimensions improve, and 3) how refiners behave. In this paper, we present a systematic study of document-level literary translation, covering nine LLMs and seven language pairs. Across nine translation-refinement granularity combinations and five refinement strategies, we find a robust recipe: document-level MT followed by segment-level refinement yields strong and stable improvements. In contrast, document-level refinement often makes fewer edits and leads to smaller or less reliable gains. Beyond granularity, A simple general refinement prompt consistently outperforms error-specific prompting and evaluate-then-refine schemes. Our large-scale human evaluation shows that refinement gains come primarily from fluency, style, and terminology, with limited and less consistent improvements in adequacy. Experiments varying model strength reveal refinement projects outputs toward the refiner's distribution rather than performing targeted error repair. These findings clarify the mechanisms and limitations of current refinement approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2605_13368
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle What Does LLM Refinement Actually Improve? A Systematic Study on Document-Level Literary Translation
Tan, Shaomu
Zhu, Dawei
Tran, Ke
Denkowski, Michael
Trenous, Sony
Byrne, Bill
Ribeiro, Leonardo
Hieber, Felix
Computation and Language
Iterative self-refinement is a simple inference-time strategy for machine translation: an LLM revises its own translation over multiple inference-time passes. Yet document-scale refinement remains poorly understood: 1) which pipelines work best, 2) what quality dimensions improve, and 3) how refiners behave. In this paper, we present a systematic study of document-level literary translation, covering nine LLMs and seven language pairs. Across nine translation-refinement granularity combinations and five refinement strategies, we find a robust recipe: document-level MT followed by segment-level refinement yields strong and stable improvements. In contrast, document-level refinement often makes fewer edits and leads to smaller or less reliable gains. Beyond granularity, A simple general refinement prompt consistently outperforms error-specific prompting and evaluate-then-refine schemes. Our large-scale human evaluation shows that refinement gains come primarily from fluency, style, and terminology, with limited and less consistent improvements in adequacy. Experiments varying model strength reveal refinement projects outputs toward the refiner's distribution rather than performing targeted error repair. These findings clarify the mechanisms and limitations of current refinement approaches.
title What Does LLM Refinement Actually Improve? A Systematic Study on Document-Level Literary Translation
topic Computation and Language
url https://arxiv.org/abs/2605.13368