Towards Explainable Evaluation Metrics for Machine Translation
Fuente:
arXiv
Saved in:
| Main Authors: | Leiter, Christoph, Lertvittayakumjorn, Piyawat, Fomicheva, Marina, Zhao, Wei, Gao, Yang, Eger, Steffen |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PrExMe! Large Scale Prompt Exploration of Open Source LLMs for Machine Translation and Summarization Evaluation
by: Leiter, Christoph, et al.
Published: (2024)
by: Leiter, Christoph, et al.
Published: (2024)
BMX: Boosting Natural Language Generation Metrics with Explainability
by: Leiter, Christoph, et al.
Published: (2022)
by: Leiter, Christoph, et al.
Published: (2022)
LLM Analysis of 150+ years of German Parliamentary Debates on Migration Reveals Shift from Post-War Solidarity to Anti-Solidarity in the Last Decade
by: Kostikova, Aida, et al.
Published: (2025)
by: Kostikova, Aida, et al.
Published: (2025)
USCORE: An Effective Approach to Fully Unsupervised Evaluation Metrics for Machine Translation
by: Belouadi, Jonas, et al.
Published: (2022)
by: Belouadi, Jonas, et al.
Published: (2022)
CROC: Evaluating and Training T2I Metrics with Pseudo- and Human-Labeled Contrastive Robustness Checks
by: Leiter, Christoph, et al.
Published: (2025)
by: Leiter, Christoph, et al.
Published: (2025)
Usable XAI: 10 Strategies Towards Exploiting Explainability in the LLM Era
by: Wu, Xuansheng, et al.
Published: (2024)
by: Wu, Xuansheng, et al.
Published: (2024)
NLLG Quarterly arXiv Report 09/24: What are the most influential current AI Papers?
by: Leiter, Christoph, et al.
Published: (2024)
by: Leiter, Christoph, et al.
Published: (2024)
BatchGEMBA: Token-Efficient Machine Translation Evaluation with Batched Prompting and Prompt Compression
by: Larionov, Daniil, et al.
Published: (2025)
by: Larionov, Daniil, et al.
Published: (2025)
GerAV: Towards New Heights in German Authorship Verification using Fine-Tuned LLMs on a New Benchmark
by: Kiefer, Lotta, et al.
Published: (2026)
by: Kiefer, Lotta, et al.
Published: (2026)
Can Capacitive Touch Images Enhance Mobile Keyboard Decoding?
by: Lertvittayakumjorn, Piyawat, et al.
Published: (2024)
by: Lertvittayakumjorn, Piyawat, et al.
Published: (2024)
AI Enabled User-Specific Cyberbullying Severity Detection with Explainability
by: Prama, Tabia Tanzin, et al.
Published: (2025)
by: Prama, Tabia Tanzin, et al.
Published: (2025)
How Good Are LLMs for Literary Translation, Really? Literary Translation Evaluation with Humans and LLMs
by: Zhang, Ran, et al.
Published: (2024)
by: Zhang, Ran, et al.
Published: (2024)
Diagnosing Hate Speech Classification: Where Do Humans and Machines Disagree, and Why?
by: Yang, Xilin
Published: (2024)
by: Yang, Xilin
Published: (2024)
ClaimVer: Explainable Claim-Level Verification and Evidence Attribution of Text Through Knowledge Graphs
by: Dammu, Preetam Prabhu Srikar, et al.
Published: (2024)
by: Dammu, Preetam Prabhu Srikar, et al.
Published: (2024)
Towards Geo-Culturally Grounded LLM Generations
by: Lertvittayakumjorn, Piyawat, et al.
Published: (2025)
by: Lertvittayakumjorn, Piyawat, et al.
Published: (2025)
Translational Gaps in Graph Transformers for Longitudinal EHR Prediction: A Critical Appraisal of GT-BEHRT
by: Tadigotla, Krish
Published: (2026)
by: Tadigotla, Krish
Published: (2026)
Subtle Biases Need Subtler Measures: Dual Metrics for Evaluating Representative and Affinity Bias in Large Language Models
by: Kumar, Abhishek, et al.
Published: (2024)
by: Kumar, Abhishek, et al.
Published: (2024)
SoK: Machine Learning for Misinformation Detection
by: Xiao, Madelyne, et al.
Published: (2023)
by: Xiao, Madelyne, et al.
Published: (2023)
PromptOptMe: Error-Aware Prompt Compression for LLM-based MT Evaluation Metrics
by: Larionov, Daniil, et al.
Published: (2024)
by: Larionov, Daniil, et al.
Published: (2024)
ELMES: An Automated Framework for Evaluating Large Language Models in Educational Scenarios
by: Wei, Shou'ang, et al.
Published: (2025)
by: Wei, Shou'ang, et al.
Published: (2025)
Towards Robust Evaluation of Unlearning in LLMs via Data Transformations
by: Joshi, Abhinav, et al.
Published: (2024)
by: Joshi, Abhinav, et al.
Published: (2024)
Identifying Climate Targets in National Laws and Policies using Machine Learning
by: Juhasz, Matyas, et al.
Published: (2024)
by: Juhasz, Matyas, et al.
Published: (2024)
Towards Modeling Learner Performance with Large Language Models
by: Neshaei, Seyed Parsa, et al.
Published: (2024)
by: Neshaei, Seyed Parsa, et al.
Published: (2024)
Toward LLM-Supported Automated Assessment of Critical Thinking Subskills
by: Peczuh, Marisa C., et al.
Published: (2025)
by: Peczuh, Marisa C., et al.
Published: (2025)
Why Don't Prompt-Based Fairness Metrics Correlate?
by: Zayed, Abdelrahman, et al.
Published: (2024)
by: Zayed, Abdelrahman, et al.
Published: (2024)
Estimating Item Difficulty Using Large Language Models and Tree-Based Machine Learning Algorithms
by: Razavi, Pooya, et al.
Published: (2025)
by: Razavi, Pooya, et al.
Published: (2025)
Evaluating the Promise and Pitfalls of LLMs in Hiring Decisions
by: Anzenberg, Eitan, et al.
Published: (2025)
by: Anzenberg, Eitan, et al.
Published: (2025)
Towards Unsupervised Question Answering System with Multi-level Summarization for Legal Text
by: Prabhu, M Manvith, et al.
Published: (2024)
by: Prabhu, M Manvith, et al.
Published: (2024)
Toward Automated Detection of Biased Social Signals from the Content of Clinical Conversations
by: Chen, Feng, et al.
Published: (2024)
by: Chen, Feng, et al.
Published: (2024)
Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is Offensive
by: Weerasooriya, Tharindu Cyril, et al.
Published: (2023)
by: Weerasooriya, Tharindu Cyril, et al.
Published: (2023)
PRSM: A Measure to Evaluate CLIP's Robustness Against Paraphrases
by: Schlegel, Udo, et al.
Published: (2025)
by: Schlegel, Udo, et al.
Published: (2025)
RTP-LX: Can LLMs Evaluate Toxicity in Multilingual Scenarios?
by: de Wynter, Adrian, et al.
Published: (2024)
by: de Wynter, Adrian, et al.
Published: (2024)
Evaluating GPT-4 at Grading Handwritten Solutions in Math Exams
by: Caraeni, Adriana, et al.
Published: (2024)
by: Caraeni, Adriana, et al.
Published: (2024)
Explainable Artificial Intelligence: A Survey of Needs, Techniques, Applications, and Future Direction
by: Mersha, Melkamu, et al.
Published: (2024)
by: Mersha, Melkamu, et al.
Published: (2024)
Beyond Overcorrection: Evaluating Diversity in T2I Models with DivBench
by: Friedrich, Felix, et al.
Published: (2025)
by: Friedrich, Felix, et al.
Published: (2025)
RDBE: Reasoning Distillation-Based Evaluation Enhances Automatic Essay Scoring
by: Mohammadkhani, Ali Ghiasvand
Published: (2024)
by: Mohammadkhani, Ali Ghiasvand
Published: (2024)
Generalization in Healthcare AI: Evaluation of a Clinical Large Language Model
by: Rahman, Salman, et al.
Published: (2024)
by: Rahman, Salman, et al.
Published: (2024)
Customize Multi-modal RAI Guardrails with Precedent-based predictions
by: Yang, Cheng-Fu, et al.
Published: (2025)
by: Yang, Cheng-Fu, et al.
Published: (2025)
Toward Cultural Interpretability: A Linguistic Anthropological Framework for Describing and Evaluating Large Language Models (LLMs)
by: Jones, Graham M., et al.
Published: (2024)
by: Jones, Graham M., et al.
Published: (2024)
Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
by: Majumdar, Ayan, et al.
Published: (2025)
by: Majumdar, Ayan, et al.
Published: (2025)
Similar Items
-
PrExMe! Large Scale Prompt Exploration of Open Source LLMs for Machine Translation and Summarization Evaluation
by: Leiter, Christoph, et al.
Published: (2024) -
BMX: Boosting Natural Language Generation Metrics with Explainability
by: Leiter, Christoph, et al.
Published: (2022) -
LLM Analysis of 150+ years of German Parliamentary Debates on Migration Reveals Shift from Post-War Solidarity to Anti-Solidarity in the Last Decade
by: Kostikova, Aida, et al.
Published: (2025) -
USCORE: An Effective Approach to Fully Unsupervised Evaluation Metrics for Machine Translation
by: Belouadi, Jonas, et al.
Published: (2022) -
CROC: Evaluating and Training T2I Metrics with Pseudo- and Human-Labeled Contrastive Robustness Checks
by: Leiter, Christoph, et al.
Published: (2025)