LLMs Are Not Scorers: Rethinking MT Evaluation with Generation-Based Methods
Fuente:
arXiv
Salvato in:
| Autore principale: | Cui, Hyang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Omnilingual MT: Machine Translation for 1,600 Languages
di: Omnilingual MT Team, et al.
Pubblicazione: (2026)
di: Omnilingual MT Team, et al.
Pubblicazione: (2026)
Homogeneous Keys, Heterogeneous Values: Exploiting Local KV Cache Asymmetry for Long-Context LLMs
di: Cui, Wanyun, et al.
Pubblicazione: (2025)
di: Cui, Wanyun, et al.
Pubblicazione: (2025)
MT-Ranker: Reference-free machine translation evaluation by inter-system ranking
di: Moosa, Ibraheem Muhammad, et al.
Pubblicazione: (2024)
di: Moosa, Ibraheem Muhammad, et al.
Pubblicazione: (2024)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
di: Saji, Alan, et al.
Pubblicazione: (2025)
di: Saji, Alan, et al.
Pubblicazione: (2025)
When Retrieval Succeeds and Fails: Rethinking Retrieval-Augmented Generation for LLMs
di: Wang, Yongjie, et al.
Pubblicazione: (2025)
di: Wang, Yongjie, et al.
Pubblicazione: (2025)
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
di: Simhi, Adi, et al.
Pubblicazione: (2025)
di: Simhi, Adi, et al.
Pubblicazione: (2025)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
di: Ashuach, Tomer, et al.
Pubblicazione: (2025)
di: Ashuach, Tomer, et al.
Pubblicazione: (2025)
Overview of the Sensemaking Task at the ELOQUENT 2025 Lab: LLMs as Teachers, Students and Evaluators
di: Šindelář, Pavel, et al.
Pubblicazione: (2025)
di: Šindelář, Pavel, et al.
Pubblicazione: (2025)
OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs
di: Wang, Yuxia, et al.
Pubblicazione: (2024)
di: Wang, Yuxia, et al.
Pubblicazione: (2024)
A Case Study of Cross-Lingual Zero-Shot Generalization for Classical Languages in LLMs
di: Akavarapu, V. S. D. S. Mahesh, et al.
Pubblicazione: (2025)
di: Akavarapu, V. S. D. S. Mahesh, et al.
Pubblicazione: (2025)
EduGuardBench: A Holistic Benchmark for Evaluating the Pedagogical Fidelity and Adversarial Safety of LLMs as Simulated Teachers
di: Jiang, Yilin, et al.
Pubblicazione: (2025)
di: Jiang, Yilin, et al.
Pubblicazione: (2025)
Meta-Evaluation of Translation Evaluation Methods: a systematic up-to-date overview
di: Han, Lifeng, et al.
Pubblicazione: (2016)
di: Han, Lifeng, et al.
Pubblicazione: (2016)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
di: Oketunji, Abiodun Finbarrs
Pubblicazione: (2023)
di: Oketunji, Abiodun Finbarrs
Pubblicazione: (2023)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
di: Peters, Sydney, et al.
Pubblicazione: (2025)
di: Peters, Sydney, et al.
Pubblicazione: (2025)
Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation
di: Dejl, Adam, et al.
Pubblicazione: (2025)
di: Dejl, Adam, et al.
Pubblicazione: (2025)
RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors
di: Dugan, Liam, et al.
Pubblicazione: (2024)
di: Dugan, Liam, et al.
Pubblicazione: (2024)
OPOR-Bench: Evaluating Large Language Models on Online Public Opinion Report Generation
di: Yu, Jinzheng, et al.
Pubblicazione: (2025)
di: Yu, Jinzheng, et al.
Pubblicazione: (2025)
Enhancing Paraphrase Type Generation: The Impact of DPO and RLHF Evaluated with Human-Ranked Data
di: Lübbers, Christopher Lee
Pubblicazione: (2025)
di: Lübbers, Christopher Lee
Pubblicazione: (2025)
SAGE: Hierarchical LLM-Based Literary Evaluation through Ontology-Grounded Interpretive Dimensions
di: Wang, Tianyu, et al.
Pubblicazione: (2026)
di: Wang, Tianyu, et al.
Pubblicazione: (2026)
Y-NQ: English-Yorùbá Evaluation dataset for Open-Book Reading Comprehension and Text Generation
di: Costa-jussà, Marta R., et al.
Pubblicazione: (2024)
di: Costa-jussà, Marta R., et al.
Pubblicazione: (2024)
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
di: Simhi, Adi, et al.
Pubblicazione: (2024)
di: Simhi, Adi, et al.
Pubblicazione: (2024)
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
di: Chehbouni, Khaoula, et al.
Pubblicazione: (2025)
di: Chehbouni, Khaoula, et al.
Pubblicazione: (2025)
MIRIAD: Augmenting LLMs with millions of medical query-response pairs
di: Zheng, Qinyue, et al.
Pubblicazione: (2025)
di: Zheng, Qinyue, et al.
Pubblicazione: (2025)
RAG-Optimized Tibetan Tourism LLMs: Enhancing Accuracy and Personalization
di: Qi, Jinhu, et al.
Pubblicazione: (2024)
di: Qi, Jinhu, et al.
Pubblicazione: (2024)
MAWARITH: A Dataset and Benchmark for Legal Inheritance Reasoning with LLMs
di: Bouchekif, Abdessalam, et al.
Pubblicazione: (2026)
di: Bouchekif, Abdessalam, et al.
Pubblicazione: (2026)
Fine-Tuning LLMs with Fine-Grained Human Feedback on Text Spans
di: CH-Wang, Sky, et al.
Pubblicazione: (2025)
di: CH-Wang, Sky, et al.
Pubblicazione: (2025)
LLMs and the Human Condition
di: Wallis, Peter
Pubblicazione: (2024)
di: Wallis, Peter
Pubblicazione: (2024)
Exploration of Augmentation Strategies in Multi-modal Retrieval-Augmented Generation for the Biomedical Domain: A Case Study Evaluating Question Answering in Glycobiology
di: Kocbek, Primož, et al.
Pubblicazione: (2025)
di: Kocbek, Primož, et al.
Pubblicazione: (2025)
Just Pass Twice: Efficient Token Classification with LLMs for Zero-Shot NER
di: Ewais, Ahmed, et al.
Pubblicazione: (2026)
di: Ewais, Ahmed, et al.
Pubblicazione: (2026)
GroUSE: A Benchmark to Evaluate Evaluators in Grounded Question Answering
di: Muller, Sacha, et al.
Pubblicazione: (2024)
di: Muller, Sacha, et al.
Pubblicazione: (2024)
MALT: Mechanistic Ablation of Lossy Translation in LLMs for a Low-Resource Language: Urdu
di: Bajwa, Taaha Saleem
Pubblicazione: (2025)
di: Bajwa, Taaha Saleem
Pubblicazione: (2025)
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
di: Simhi, Adi, et al.
Pubblicazione: (2025)
di: Simhi, Adi, et al.
Pubblicazione: (2025)
Temporal Referential Consistency: Do LLMs Favor Sequences Over Absolute Time References?
di: Bajpai, Ashutosh, et al.
Pubblicazione: (2025)
di: Bajpai, Ashutosh, et al.
Pubblicazione: (2025)
Improving the OOD Performance of Closed-Source LLMs on NLI Through Strategic Data Selection
di: Stacey, Joe, et al.
Pubblicazione: (2025)
di: Stacey, Joe, et al.
Pubblicazione: (2025)
LLMs Generate Kitsch
di: Klinge, Xenia, et al.
Pubblicazione: (2026)
di: Klinge, Xenia, et al.
Pubblicazione: (2026)
Cetvel: A Unified Benchmark for Evaluating Language Understanding, Generation and Cultural Capacity of LLMs for Turkish
di: Er, Yakup Abrek, et al.
Pubblicazione: (2025)
di: Er, Yakup Abrek, et al.
Pubblicazione: (2025)
All for One: LLMs Solve Mental Math at the Last Token With Information Transferred From Other Tokens
di: Mamidanna, Siddarth, et al.
Pubblicazione: (2025)
di: Mamidanna, Siddarth, et al.
Pubblicazione: (2025)
Improving LLMs with a knowledge from databases
di: Máša, Petr
Pubblicazione: (2025)
di: Máša, Petr
Pubblicazione: (2025)
Assessing RAG and HyDE on 1B vs. 4B-Parameter Gemma LLMs for Personal Assistants Integretion
di: Sorstkins, Andrejs
Pubblicazione: (2025)
di: Sorstkins, Andrejs
Pubblicazione: (2025)
Layer-Aware Embedding Fusion for LLMs in Text Classifications
di: Gwak, Jiho, et al.
Pubblicazione: (2025)
di: Gwak, Jiho, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Omnilingual MT: Machine Translation for 1,600 Languages
di: Omnilingual MT Team, et al.
Pubblicazione: (2026) -
Homogeneous Keys, Heterogeneous Values: Exploiting Local KV Cache Asymmetry for Long-Context LLMs
di: Cui, Wanyun, et al.
Pubblicazione: (2025) -
MT-Ranker: Reference-free machine translation evaluation by inter-system ranking
di: Moosa, Ibraheem Muhammad, et al.
Pubblicazione: (2024) -
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
di: Saji, Alan, et al.
Pubblicazione: (2025) -
When Retrieval Succeeds and Fails: Rethinking Retrieval-Augmented Generation for LLMs
di: Wang, Yongjie, et al.
Pubblicazione: (2025)