Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation
Fuente:
arXiv
Saved in:
| Main Authors: | Kocyigit, Muhammed Yusuf, Briakou, Eleftheria, Deutsch, Daniel, Luo, Jiaming, Cherry, Colin, Freitag, Markus |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Translating Step-by-Step: Decomposing the Translation Process for Improved Translation Quality of Long-Form Texts
by: Briakou, Eleftheria, et al.
Published: (2024)
by: Briakou, Eleftheria, et al.
Published: (2024)
On the Implications of Verbose LLM Outputs: A Case Study in Translation Evaluation
by: Briakou, Eleftheria, et al.
Published: (2024)
by: Briakou, Eleftheria, et al.
Published: (2024)
Leveraging Domain Knowledge at Inference Time for LLM Translation: Retrieval versus Generation
by: Li, Bryan, et al.
Published: (2025)
by: Li, Bryan, et al.
Published: (2025)
The Impact of Post-training on Data Contamination
by: Kocyigit, Muhammed Yusuf, et al.
Published: (2026)
by: Kocyigit, Muhammed Yusuf, et al.
Published: (2026)
Déjà Vu: Multilingual LLM Evaluation through the Lens of Machine Translation Evaluation
by: Kreutzer, Julia, et al.
Published: (2025)
by: Kreutzer, Julia, et al.
Published: (2025)
Enhancing Human Evaluation in Machine Translation with Comparative Judgment
by: Song, Yixiao, et al.
Published: (2025)
by: Song, Yixiao, et al.
Published: (2025)
To Diverge or Not to Diverge: A Morphosyntactic Perspective on Machine Translation vs Human Translation
by: Luo, Jiaming, et al.
Published: (2024)
by: Luo, Jiaming, et al.
Published: (2024)
Don't Throw Away Data: Better Sequence Knowledge Distillation
by: Wang, Jun, et al.
Published: (2024)
by: Wang, Jun, et al.
Published: (2024)
TranslateGemma Technical Report
by: Finkelstein, Mara, et al.
Published: (2026)
by: Finkelstein, Mara, et al.
Published: (2026)
Beyond Human-Only: Evaluating Human-Machine Collaboration for Collecting High-Quality Translation Data
by: Liu, Zhongtao, et al.
Published: (2024)
by: Liu, Zhongtao, et al.
Published: (2024)
MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation
by: Riley, Parker, et al.
Published: (2025)
by: Riley, Parker, et al.
Published: (2025)
SSA-COMET: Do LLMs Outperform Learned Metrics in Evaluating MT for Under-Resourced African Languages?
by: Li, Senyu, et al.
Published: (2025)
by: Li, Senyu, et al.
Published: (2025)
Mitigating Metric Bias in Minimum Bayes Risk Decoding
by: Kovacs, Geza, et al.
Published: (2024)
by: Kovacs, Geza, et al.
Published: (2024)
Rethinking Cross-lingual Alignment: Balancing Transfer and Cultural Erasure in Multilingual LLMs
by: Han, HyoJung, et al.
Published: (2025)
by: Han, HyoJung, et al.
Published: (2025)
Quality-Aware Translation Models: Efficient Generation and Quality Estimation in a Single Model
by: Tomani, Christian, et al.
Published: (2023)
by: Tomani, Christian, et al.
Published: (2023)
Generating Difficult-to-Translate Texts
by: Zouhar, Vilém, et al.
Published: (2025)
by: Zouhar, Vilém, et al.
Published: (2025)
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
by: Xu, Wenda, et al.
Published: (2025)
by: Xu, Wenda, et al.
Published: (2025)
When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method
by: Zhang, Biao, et al.
Published: (2024)
by: Zhang, Biao, et al.
Published: (2024)
MetricX-24: The Google Submission to the WMT 2024 Metrics Shared Task
by: Juraska, Juraj, et al.
Published: (2024)
by: Juraska, Juraj, et al.
Published: (2024)
MetricX-25 and GemSpanEval: Google Translate Submissions to the WMT25 Evaluation Shared Task
by: Juraska, Juraj, et al.
Published: (2025)
by: Juraska, Juraj, et al.
Published: (2025)
Finding Replicable Human Evaluations via Stable Ranking Probability
by: Riley, Parker, et al.
Published: (2024)
by: Riley, Parker, et al.
Published: (2024)
Investigating the Impact of Data Contamination of Large Language Models in Text-to-SQL Translation
by: Ranaldi, Federico, et al.
Published: (2024)
by: Ranaldi, Federico, et al.
Published: (2024)
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set
by: Finkelstein, Mara, et al.
Published: (2024)
by: Finkelstein, Mara, et al.
Published: (2024)
Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?
by: Singh, Aaditya K., et al.
Published: (2024)
by: Singh, Aaditya K., et al.
Published: (2024)
WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects
by: Deutsch, Daniel, et al.
Published: (2025)
by: Deutsch, Daniel, et al.
Published: (2025)
Feeding Two Birds or Favoring One? Adequacy-Fluency Tradeoffs in Evaluation and Meta-Evaluation of Machine Translation
by: Shayegh, Behzad, et al.
Published: (2025)
by: Shayegh, Behzad, et al.
Published: (2025)
When Flores Bloomz Wrong: Cross-Direction Contamination in Machine Translation Evaluation
by: Tan, David, et al.
Published: (2026)
by: Tan, David, et al.
Published: (2026)
Data Contamination in Neural Hieroglyphic Translation: A Reproducibility Study
by: Toutou, Ammar, et al.
Published: (2026)
by: Toutou, Ammar, et al.
Published: (2026)
Introducing the NewsPaLM MBR and QE Dataset: LLM-Generated High-Quality Parallel Data Outperforms Traditional Web-Crawled Data
by: Finkelstein, Mara, et al.
Published: (2024)
by: Finkelstein, Mara, et al.
Published: (2024)
Mind the Gap... or Not? How Translation Errors and Evaluation Details Skew Multilingual Results
by: Peter, Jan-Thorsten, et al.
Published: (2025)
by: Peter, Jan-Thorsten, et al.
Published: (2025)
Evaluating LLM-Contaminated Crowdsourcing Data Without Ground Truth
by: Zhang, Yichi, et al.
Published: (2025)
by: Zhang, Yichi, et al.
Published: (2025)
RADAR: Mechanistic Pathways for Detecting Data Contamination in LLM Evaluation
by: Kattamuri, Ashish, et al.
Published: (2025)
by: Kattamuri, Ashish, et al.
Published: (2025)
A Generic Machine Learning Framework for Fully-Unsupervised Anomaly Detection with Contaminated Data
by: Ulmer, Markus, et al.
Published: (2023)
by: Ulmer, Markus, et al.
Published: (2023)
LLMRefine: Pinpointing and Refining Large Language Models via Fine-Grained Actionable Feedback
by: Xu, Wenda, et al.
Published: (2023)
by: Xu, Wenda, et al.
Published: (2023)
Scaling Model and Data for Multilingual Machine Translation with Open Large Language Models
by: Shang, Yuzhe, et al.
Published: (2026)
by: Shang, Yuzhe, et al.
Published: (2026)
Finding Inter-species Associations on Large Citizen Science Datasets
by: Deutsch, Jacob
Published: (2025)
by: Deutsch, Jacob
Published: (2025)
Investigation of Glandula Uropygialis in Different Avian Species Using Morphometric and Histological Methods
by: Funda Aksunger Karaavci, et al.
Published: (2024)
by: Funda Aksunger Karaavci, et al.
Published: (2024)
Regularized Overestimated Newton
by: Duan, Danny, et al.
Published: (2025)
by: Duan, Danny, et al.
Published: (2025)
Evaluation of a New Non‐Mydriatic Handheld Fundus Camera for Fundus Imaging in Cats: A Retrospective Study: 208 Cases (2023–2024)
by: Özlem Şengöz Şirin, et al.
Published: (2025)
by: Özlem Şengöz Şirin, et al.
Published: (2025)
VeriContaminated: Assessing LLM-Driven Verilog Coding for Data Contamination
by: Wang, Zeng, et al.
Published: (2025)
by: Wang, Zeng, et al.
Published: (2025)
Similar Items
-
Translating Step-by-Step: Decomposing the Translation Process for Improved Translation Quality of Long-Form Texts
by: Briakou, Eleftheria, et al.
Published: (2024) -
On the Implications of Verbose LLM Outputs: A Case Study in Translation Evaluation
by: Briakou, Eleftheria, et al.
Published: (2024) -
Leveraging Domain Knowledge at Inference Time for LLM Translation: Retrieval versus Generation
by: Li, Bryan, et al.
Published: (2025) -
The Impact of Post-training on Data Contamination
by: Kocyigit, Muhammed Yusuf, et al.
Published: (2026) -
Déjà Vu: Multilingual LLM Evaluation through the Lens of Machine Translation Evaluation
by: Kreutzer, Julia, et al.
Published: (2025)