How Important is `Perfect' English for Machine Translation Prompts?

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Schmidtová, Patrícia, Bafna, Niyati, Aycock, Seth, Vico, Gianluca, Kamzela, Wiktor, Hämmerl, Katharina, Zouhar, Vilém
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914014938791936
author Schmidtová, Patrícia
Bafna, Niyati
Aycock, Seth
Vico, Gianluca
Kamzela, Wiktor
Hämmerl, Katharina
Zouhar, Vilém
author_facet Schmidtová, Patrícia
Bafna, Niyati
Aycock, Seth
Vico, Gianluca
Kamzela, Wiktor
Hämmerl, Katharina
Zouhar, Vilém
contents Large language models (LLMs) have achieved top results in recent machine translation evaluations, but they are also known to be sensitive to errors and perturbations in their prompts. We systematically evaluate how both humanly plausible and synthetic errors in user prompts affect LLMs' performance on two related tasks: Machine translation and machine translation evaluation. We provide both a quantitative analysis and qualitative insights into how the models respond to increasing noise in the user prompt. The prompt quality strongly affects the translation performance: With many errors, even a good prompt can underperform a minimal or poor prompt without errors. However, different noise types impact translation quality differently, with character-level and combined noisers degrading performance more than phrasal perturbations. Qualitative analysis reveals that lower prompt quality largely leads to poorer instruction following, rather than directly affecting translation quality itself. Further, LLMs can still translate in scenarios with overwhelming random noise that would make the prompt illegible to humans.
format Preprint
id arxiv_https___arxiv_org_abs_2507_09509
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle How Important is `Perfect' English for Machine Translation Prompts?
Schmidtová, Patrícia
Bafna, Niyati
Aycock, Seth
Vico, Gianluca
Kamzela, Wiktor
Hämmerl, Katharina
Zouhar, Vilém
Computation and Language
Large language models (LLMs) have achieved top results in recent machine translation evaluations, but they are also known to be sensitive to errors and perturbations in their prompts. We systematically evaluate how both humanly plausible and synthetic errors in user prompts affect LLMs' performance on two related tasks: Machine translation and machine translation evaluation. We provide both a quantitative analysis and qualitative insights into how the models respond to increasing noise in the user prompt. The prompt quality strongly affects the translation performance: With many errors, even a good prompt can underperform a minimal or poor prompt without errors. However, different noise types impact translation quality differently, with character-level and combined noisers degrading performance more than phrasal perturbations. Qualitative analysis reveals that lower prompt quality largely leads to poorer instruction following, rather than directly affecting translation quality itself. Further, LLMs can still translate in scenarios with overwhelming random noise that would make the prompt illegible to humans.
title How Important is `Perfect' English for Machine Translation Prompts?
topic Computation and Language
url https://arxiv.org/abs/2507.09509