Déjà Vu: Multilingual LLM Evaluation through the Lens of Machine Translation Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Kreutzer, Julia, Briakou, Eleftheria, Agrawal, Sweta, Fadaee, Marzieh, Tom, Kocmi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking Cross-lingual Alignment: Balancing Transfer and Cultural Erasure in Multilingual LLMs
by: Han, HyoJung, et al.
Published: (2025)
by: Han, HyoJung, et al.
Published: (2025)
Unlocking Reasoning Capability on Machine Translation in Large Language Models
by: Rajaee, Sara, et al.
Published: (2026)
by: Rajaee, Sara, et al.
Published: (2026)
Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation
by: Kocyigit, Muhammed Yusuf, et al.
Published: (2025)
by: Kocyigit, Muhammed Yusuf, et al.
Published: (2025)
SLIDE: Reference-free Evaluation for Machine Translation using a Sliding Document Window
by: Raunak, Vikas, et al.
Published: (2023)
by: Raunak, Vikas, et al.
Published: (2023)
LLM See, LLM Do: Guiding Data Generation to Target Non-Differentiable Objectives
by: Shimabucoro, Luísa, et al.
Published: (2024)
by: Shimabucoro, Luísa, et al.
Published: (2024)
The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm
by: Aakanksha, et al.
Published: (2024)
by: Aakanksha, et al.
Published: (2024)
A Post-trainer's Guide to Multilingual Training Data: Uncovering Cross-lingual Transfer Dynamics
by: Shimabucoro, Luisa, et al.
Published: (2025)
by: Shimabucoro, Luisa, et al.
Published: (2025)
On the Implications of Verbose LLM Outputs: A Case Study in Translation Evaluation
by: Briakou, Eleftheria, et al.
Published: (2024)
by: Briakou, Eleftheria, et al.
Published: (2024)
A Linguistic Analysis of Spontaneous Thoughts: Investigating Experiences of Déjà Vu, Unexpected Thoughts, and Involuntary Autobiographical Memories
by: Venkatesha, Videep, et al.
Published: (2025)
by: Venkatesha, Videep, et al.
Published: (2025)
The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It
by: Yong, Zheng-Xin, et al.
Published: (2025)
by: Yong, Zheng-Xin, et al.
Published: (2025)
AI-Assisted Human Evaluation of Machine Translation
by: Zouhar, Vilém, et al.
Published: (2024)
by: Zouhar, Vilém, et al.
Published: (2024)
M-RewardBench: Evaluating Reward Models in Multilingual Settings
by: Gureja, Srishti, et al.
Published: (2024)
by: Gureja, Srishti, et al.
Published: (2024)
NeoBabel: A Multilingual Open Tower for Visual Generation
by: Derakhshani, Mohammad Mahdi, et al.
Published: (2025)
by: Derakhshani, Mohammad Mahdi, et al.
Published: (2025)
On the Evaluation Practices in Multilingual NLP: Can Machine Translation Offer an Alternative to Human Translations?
by: Choenni, Rochelle, et al.
Published: (2024)
by: Choenni, Rochelle, et al.
Published: (2024)
The Multilingual Divide and Its Impact on Global AI Safety
by: Peppin, Aidan, et al.
Published: (2025)
by: Peppin, Aidan, et al.
Published: (2025)
TranslateGemma Technical Report
by: Finkelstein, Mara, et al.
Published: (2026)
by: Finkelstein, Mara, et al.
Published: (2026)
Pearmut: Human Evaluation of Translation Made Trivial
by: Zouhar, Vilém, et al.
Published: (2026)
by: Zouhar, Vilém, et al.
Published: (2026)
MTQ-Eval: Multilingual Text Quality Evaluation for Language Models
by: Pokharel, Rhitabrat, et al.
Published: (2025)
by: Pokharel, Rhitabrat, et al.
Published: (2025)
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
by: Xu, Wenda, et al.
Published: (2025)
by: Xu, Wenda, et al.
Published: (2025)
Diversify and Conquer: Diversity-Centric Data Selection with Iterative Refinement
by: Yu, Simon, et al.
Published: (2024)
by: Yu, Simon, et al.
Published: (2024)
Translation as a Scalable Proxy for Multilingual Evaluation
by: Issaka, Sheriff, et al.
Published: (2026)
by: Issaka, Sheriff, et al.
Published: (2026)
The Art of Asking: Multilingual Prompt Optimization for Synthetic Data
by: Mora, David, et al.
Published: (2025)
by: Mora, David, et al.
Published: (2025)
Verification Limits Code LLM Training
by: Gureja, Srishti, et al.
Published: (2025)
by: Gureja, Srishti, et al.
Published: (2025)
PMMT: Preference Alignment in Multilingual Machine Translation via LLM Distillation
by: Sun, Shuqiao, et al.
Published: (2024)
by: Sun, Shuqiao, et al.
Published: (2024)
Leveraging Domain Knowledge at Inference Time for LLM Translation: Retrieval versus Generation
by: Li, Bryan, et al.
Published: (2025)
by: Li, Bryan, et al.
Published: (2025)
Translate Smart, not Hard: Cascaded Translation Systems with Quality-Aware Deferral
by: Farinhas, António, et al.
Published: (2025)
by: Farinhas, António, et al.
Published: (2025)
Machine Translation Meta Evaluation through Translation Accuracy Challenge Sets
by: Moghe, Nikita, et al.
Published: (2024)
by: Moghe, Nikita, et al.
Published: (2024)
SSA-COMET: Do LLMs Outperform Learned Metrics in Evaluating MT for Under-Resourced African Languages?
by: Li, Senyu, et al.
Published: (2025)
by: Li, Senyu, et al.
Published: (2025)
Making, not Taking, the Best of N
by: Khairi, Ammar, et al.
Published: (2025)
by: Khairi, Ammar, et al.
Published: (2025)
When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs
by: Khairi, Ammar, et al.
Published: (2025)
by: Khairi, Ammar, et al.
Published: (2025)
Translating Step-by-Step: Decomposing the Translation Process for Improved Translation Quality of Long-Form Texts
by: Briakou, Eleftheria, et al.
Published: (2024)
by: Briakou, Eleftheria, et al.
Published: (2024)
News Deja Vu: Connecting Past and Present with Semantic Search
by: Franklin, Brevin, et al.
Published: (2024)
by: Franklin, Brevin, et al.
Published: (2024)
On the Shortcut Learning in Multilingual Neural Machine Translation
by: Wang, Wenxuan, et al.
Published: (2024)
by: Wang, Wenxuan, et al.
Published: (2024)
Transformer-Encoder Trees for Efficient Multilingual Machine Translation and Speech Translation
by: Guan, Yiwen, et al.
Published: (2025)
by: Guan, Yiwen, et al.
Published: (2025)
Span-Level Machine Translation Meta-Evaluation
by: Perrella, Stefano, et al.
Published: (2026)
by: Perrella, Stefano, et al.
Published: (2026)
DuoLens: A Framework for Robust Detection of Machine-Generated Multilingual Text and Code
by: Agrawal, Shriyansh, et al.
Published: (2025)
by: Agrawal, Shriyansh, et al.
Published: (2025)
Beyond Correlation: Interpretable Evaluation of Machine Translation Metrics
by: Perrella, Stefano, et al.
Published: (2024)
by: Perrella, Stefano, et al.
Published: (2024)
Evaluation of Machine Translation Based on Semantic Dependencies and Keywords
by: Yuan, Kewei, et al.
Published: (2024)
by: Yuan, Kewei, et al.
Published: (2024)
Code-Switching Red-Teaming: LLM Evaluation for Safety and Multilingual Understanding
by: Yoo, Haneul, et al.
Published: (2024)
by: Yoo, Haneul, et al.
Published: (2024)
Towards Multilingual LLM Evaluation for European Languages
by: Thellmann, Klaudia, et al.
Published: (2024)
by: Thellmann, Klaudia, et al.
Published: (2024)
Similar Items
-
Rethinking Cross-lingual Alignment: Balancing Transfer and Cultural Erasure in Multilingual LLMs
by: Han, HyoJung, et al.
Published: (2025) -
Unlocking Reasoning Capability on Machine Translation in Large Language Models
by: Rajaee, Sara, et al.
Published: (2026) -
Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation
by: Kocyigit, Muhammed Yusuf, et al.
Published: (2025) -
SLIDE: Reference-free Evaluation for Machine Translation using a Sliding Document Window
by: Raunak, Vikas, et al.
Published: (2023) -
LLM See, LLM Do: Guiding Data Generation to Target Non-Differentiable Objectives
by: Shimabucoro, Luísa, et al.
Published: (2024)