Generating Difficult-to-Translate Texts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zouhar, Vilém, Xu, Wenda, Riley, Parker, Juraska, Juraj, Finkelstein, Mara, Freitag, Markus, Deutsch, Daniel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918152882880512
author Zouhar, Vilém
Xu, Wenda
Riley, Parker
Juraska, Juraj
Finkelstein, Mara
Freitag, Markus
Deutsch, Daniel
author_facet Zouhar, Vilém
Xu, Wenda
Riley, Parker
Juraska, Juraj
Finkelstein, Mara
Freitag, Markus
Deutsch, Daniel
contents Machine translation benchmarks sourced from the real world are quickly obsoleted, due to most examples being easy for state-of-the-art translation models. This limits the benchmark's ability to distinguish which model is better or to reveal models' weaknesses. Current methods for creating difficult test cases, such as subsampling or from-scratch synthesis, either fall short of identifying difficult examples or suffer from a lack of diversity and naturalness. Inspired by the iterative process of human experts probing for model failures, we propose MT-breaker, a method where a large language model iteratively refines a source text to increase its translation difficulty. The LLM iteratively queries a target machine translation model to guide its generation of difficult examples. Our approach generates examples that are more challenging for the target MT model while preserving the diversity of natural texts. While the examples are tailored to a particular machine translation model during the generation, the difficulty also transfers to other models and languages.
format Preprint
id arxiv_https___arxiv_org_abs_2509_26592
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generating Difficult-to-Translate Texts
Zouhar, Vilém
Xu, Wenda
Riley, Parker
Juraska, Juraj
Finkelstein, Mara
Freitag, Markus
Deutsch, Daniel
Computation and Language
Machine translation benchmarks sourced from the real world are quickly obsoleted, due to most examples being easy for state-of-the-art translation models. This limits the benchmark's ability to distinguish which model is better or to reveal models' weaknesses. Current methods for creating difficult test cases, such as subsampling or from-scratch synthesis, either fall short of identifying difficult examples or suffer from a lack of diversity and naturalness. Inspired by the iterative process of human experts probing for model failures, we propose MT-breaker, a method where a large language model iteratively refines a source text to increase its translation difficulty. The LLM iteratively queries a target machine translation model to guide its generation of difficult examples. Our approach generates examples that are more challenging for the target MT model while preserving the diversity of natural texts. While the examples are tailored to a particular machine translation model during the generation, the difficulty also transfers to other models and languages.
title Generating Difficult-to-Translate Texts
topic Computation and Language
url https://arxiv.org/abs/2509.26592