ContrastScore: Towards Higher Quality, Less Biased, More Efficient Evaluation Metrics with Contrastive Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Xiao, Larionov, Daniil, Wu, Siwei, Liu, Yiqi, Eger, Steffen, Moosavi, Nafise Sadat, Lin, Chenghua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores
von: Liu, Yiqi, et al.
Veröffentlicht: (2023)
von: Liu, Yiqi, et al.
Veröffentlicht: (2023)
PromptOptMe: Error-Aware Prompt Compression for LLM-based MT Evaluation Metrics
von: Larionov, Daniil, et al.
Veröffentlicht: (2024)
von: Larionov, Daniil, et al.
Veröffentlicht: (2024)
BatchGEMBA: Token-Efficient Machine Translation Evaluation with Batched Prompting and Prompt Compression
von: Larionov, Daniil, et al.
Veröffentlicht: (2025)
von: Larionov, Daniil, et al.
Veröffentlicht: (2025)
More or Less Wrong: A Benchmark for Directional Bias in LLM Comparative Reasoning
von: Shafiei, Mohammadamin, et al.
Veröffentlicht: (2025)
von: Shafiei, Mohammadamin, et al.
Veröffentlicht: (2025)
RIGOURATE: Quantifying Scientific Exaggeration with Evidence-Aligned Claim Evaluation
von: James, Joseph, et al.
Veröffentlicht: (2026)
von: James, Joseph, et al.
Veröffentlicht: (2026)
xCOMET-lite: Bridging the Gap Between Efficiency and Quality in Learned MT Evaluation Metrics
von: Larionov, Daniil, et al.
Veröffentlicht: (2024)
von: Larionov, Daniil, et al.
Veröffentlicht: (2024)
Initialisation Determines the Basin: Efficient Codebook Optimisation for Extreme LLM Quantization
von: Kennedy, Ian W., et al.
Veröffentlicht: (2026)
von: Kennedy, Ian W., et al.
Veröffentlicht: (2026)
CROC: Evaluating and Training T2I Metrics with Pseudo- and Human-Labeled Contrastive Robustness Checks
von: Leiter, Christoph, et al.
Veröffentlicht: (2025)
von: Leiter, Christoph, et al.
Veröffentlicht: (2025)
How to Leverage Digit Embeddings to Represent Numbers?
von: Sivakumar, Jasivan Alex, et al.
Veröffentlicht: (2024)
von: Sivakumar, Jasivan Alex, et al.
Veröffentlicht: (2024)
USCORE: An Effective Approach to Fully Unsupervised Evaluation Metrics for Machine Translation
von: Belouadi, Jonas, et al.
Veröffentlicht: (2022)
von: Belouadi, Jonas, et al.
Veröffentlicht: (2022)
MultiHoax: A Dataset of Multi-hop False-Premise Questions
von: Shafiei, Mohammadamin, et al.
Veröffentlicht: (2025)
von: Shafiei, Mohammadamin, et al.
Veröffentlicht: (2025)
Rolling the DICE on Idiomaticity: How LLMs Fail to Grasp Context
von: Mi, Maggie, et al.
Veröffentlicht: (2024)
von: Mi, Maggie, et al.
Veröffentlicht: (2024)
From Input Perception to Predictive Insight: Modeling Model Blind Spots Before They Become Errors
von: Mi, Maggie, et al.
Veröffentlicht: (2025)
von: Mi, Maggie, et al.
Veröffentlicht: (2025)
Deconstructing Attention: Investigating Design Principles for Effective Language Modeling
von: Xue, Huiyin, et al.
Veröffentlicht: (2025)
von: Xue, Huiyin, et al.
Veröffentlicht: (2025)
Exploring the Influence of Label Aggregation on Minority Voices: Implications for Dataset Bias and Model Training
von: Pandya, Mugdha, et al.
Veröffentlicht: (2024)
von: Pandya, Mugdha, et al.
Veröffentlicht: (2024)
Hidden Failures in Robustness: Why Supervised Uncertainty Quantification Needs Better Evaluation
von: Stacey, Joe, et al.
Veröffentlicht: (2026)
von: Stacey, Joe, et al.
Veröffentlicht: (2026)
Transforming Science with Large Language Models: A Survey on AI-assisted Scientific Discovery, Experimentation, Content Generation, and Evaluation
von: Eger, Steffen, et al.
Veröffentlicht: (2025)
von: Eger, Steffen, et al.
Veröffentlicht: (2025)
Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics
von: Roy, Subhadeep, et al.
Veröffentlicht: (2026)
von: Roy, Subhadeep, et al.
Veröffentlicht: (2026)
LIME: Less Is More for MLLM Evaluation
von: Zhu, King, et al.
Veröffentlicht: (2024)
von: Zhu, King, et al.
Veröffentlicht: (2024)
Towards Explainable Evaluation Metrics for Machine Translation
von: Leiter, Christoph, et al.
Veröffentlicht: (2023)
von: Leiter, Christoph, et al.
Veröffentlicht: (2023)
Decoding News Narratives: A Critical Analysis of Large Language Models in Framing Detection
von: Pastorino, Valeria, et al.
Veröffentlicht: (2024)
von: Pastorino, Valeria, et al.
Veröffentlicht: (2024)
Less is More: Towards Simple Graph Contrastive Learning
von: Zhao, Yanan, et al.
Veröffentlicht: (2025)
von: Zhao, Yanan, et al.
Veröffentlicht: (2025)
No Shortcuts to Culture: Indonesian Multi-hop Question Answering for Complex Cultural Understanding
von: Permadi, Vynska Amalia, et al.
Veröffentlicht: (2026)
von: Permadi, Vynska Amalia, et al.
Veröffentlicht: (2026)
LLMs Do Not See Age: Assessing Demographic Bias in Automated Systematic Review Synthesis
von: Aghaebe, Favour Yahdii, et al.
Veröffentlicht: (2025)
von: Aghaebe, Favour Yahdii, et al.
Veröffentlicht: (2025)
Faithful Summarisation under Disagreement via Belief-Level Aggregation
von: Aghaebe, Favour Yahdii, et al.
Veröffentlicht: (2026)
von: Aghaebe, Favour Yahdii, et al.
Veröffentlicht: (2026)
Exploring Gender Disparities in Automatic Speech Recognition Technology
von: ElGhazaly, Hend, et al.
Veröffentlicht: (2025)
von: ElGhazaly, Hend, et al.
Veröffentlicht: (2025)
DeepSeek-R1 vs. o3-mini: How Well can Reasoning LLMs Evaluate MT and Summarization?
von: Larionov, Daniil, et al.
Veröffentlicht: (2025)
von: Larionov, Daniil, et al.
Veröffentlicht: (2025)
MMTE: Corpus and Metrics for Evaluating Machine Translation Quality of Metaphorical Language
von: Wang, Shun, et al.
Veröffentlicht: (2024)
von: Wang, Shun, et al.
Veröffentlicht: (2024)
PrExMe! Large Scale Prompt Exploration of Open Source LLMs for Machine Translation and Summarization Evaluation
von: Leiter, Christoph, et al.
Veröffentlicht: (2024)
von: Leiter, Christoph, et al.
Veröffentlicht: (2024)
Beyond Hate Speech: NLP's Challenges and Opportunities in Uncovering Dehumanizing Language
von: Saffari, Hamidreza, et al.
Veröffentlicht: (2024)
von: Saffari, Hamidreza, et al.
Veröffentlicht: (2024)
See More, Change Less: Anatomy-Aware Diffusion for Contrast Enhancement
von: Liu, Junqi, et al.
Veröffentlicht: (2025)
von: Liu, Junqi, et al.
Veröffentlicht: (2025)
TikZilla: Scaling Text-to-TikZ with High-Quality Data and Reinforcement Learning
von: Greisinger, Christian, et al.
Veröffentlicht: (2026)
von: Greisinger, Christian, et al.
Veröffentlicht: (2026)
CASPR: Automated Evaluation Metric for Contrastive Summarization
von: Ananthamurugan, Nirupan, et al.
Veröffentlicht: (2024)
von: Ananthamurugan, Nirupan, et al.
Veröffentlicht: (2024)
BMX: Boosting Natural Language Generation Metrics with Explainability
von: Leiter, Christoph, et al.
Veröffentlicht: (2022)
von: Leiter, Christoph, et al.
Veröffentlicht: (2022)
Cross-lingual Cross-temporal Summarization: Dataset, Models, Evaluation
von: Zhang, Ran, et al.
Veröffentlicht: (2023)
von: Zhang, Ran, et al.
Veröffentlicht: (2023)
How Good Are LLMs for Literary Translation, Really? Literary Translation Evaluation with Humans and LLMs
von: Zhang, Ran, et al.
Veröffentlicht: (2024)
von: Zhang, Ran, et al.
Veröffentlicht: (2024)
NLLG Quarterly arXiv Report 09/24: What are the most influential current AI Papers?
von: Leiter, Christoph, et al.
Veröffentlicht: (2024)
von: Leiter, Christoph, et al.
Veröffentlicht: (2024)
Evaluating Diversity in Automatic Poetry Generation
von: Chen, Yanran, et al.
Veröffentlicht: (2024)
von: Chen, Yanran, et al.
Veröffentlicht: (2024)
Topological Metric for Unsupervised Embedding Quality Evaluation
von: Shestov, Aleksei, et al.
Veröffentlicht: (2025)
von: Shestov, Aleksei, et al.
Veröffentlicht: (2025)
Beyond One-Size-Fits-All: Inversion Learning for Highly Effective NLG Evaluation Prompts
von: Hong, Hanhua, et al.
Veröffentlicht: (2025)
von: Hong, Hanhua, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores
von: Liu, Yiqi, et al.
Veröffentlicht: (2023) -
PromptOptMe: Error-Aware Prompt Compression for LLM-based MT Evaluation Metrics
von: Larionov, Daniil, et al.
Veröffentlicht: (2024) -
BatchGEMBA: Token-Efficient Machine Translation Evaluation with Batched Prompting and Prompt Compression
von: Larionov, Daniil, et al.
Veröffentlicht: (2025) -
More or Less Wrong: A Benchmark for Directional Bias in LLM Comparative Reasoning
von: Shafiei, Mohammadamin, et al.
Veröffentlicht: (2025) -
RIGOURATE: Quantifying Scientific Exaggeration with Evidence-Aligned Claim Evaluation
von: James, Joseph, et al.
Veröffentlicht: (2026)