LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yiqi, Moosavi, Nafise Sadat, Lin, Chenghua |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ContrastScore: Towards Higher Quality, Less Biased, More Efficient Evaluation Metrics with Contrastive Evaluation
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
RIGOURATE: Quantifying Scientific Exaggeration with Evidence-Aligned Claim Evaluation
by: James, Joseph, et al.
Published: (2026)
by: James, Joseph, et al.
Published: (2026)
Rolling the DICE on Idiomaticity: How LLMs Fail to Grasp Context
by: Mi, Maggie, et al.
Published: (2024)
by: Mi, Maggie, et al.
Published: (2024)
How to Leverage Digit Embeddings to Represent Numbers?
by: Sivakumar, Jasivan Alex, et al.
Published: (2024)
by: Sivakumar, Jasivan Alex, et al.
Published: (2024)
Initialisation Determines the Basin: Efficient Codebook Optimisation for Extreme LLM Quantization
by: Kennedy, Ian W., et al.
Published: (2026)
by: Kennedy, Ian W., et al.
Published: (2026)
MultiHoax: A Dataset of Multi-hop False-Premise Questions
by: Shafiei, Mohammadamin, et al.
Published: (2025)
by: Shafiei, Mohammadamin, et al.
Published: (2025)
From Input Perception to Predictive Insight: Modeling Model Blind Spots Before They Become Errors
by: Mi, Maggie, et al.
Published: (2025)
by: Mi, Maggie, et al.
Published: (2025)
More or Less Wrong: A Benchmark for Directional Bias in LLM Comparative Reasoning
by: Shafiei, Mohammadamin, et al.
Published: (2025)
by: Shafiei, Mohammadamin, et al.
Published: (2025)
Exploring the Influence of Label Aggregation on Minority Voices: Implications for Dataset Bias and Model Training
by: Pandya, Mugdha, et al.
Published: (2024)
by: Pandya, Mugdha, et al.
Published: (2024)
LLMs Do Not See Age: Assessing Demographic Bias in Automated Systematic Review Synthesis
by: Aghaebe, Favour Yahdii, et al.
Published: (2025)
by: Aghaebe, Favour Yahdii, et al.
Published: (2025)
Deconstructing Attention: Investigating Design Principles for Effective Language Modeling
by: Xue, Huiyin, et al.
Published: (2025)
by: Xue, Huiyin, et al.
Published: (2025)
Decoding News Narratives: A Critical Analysis of Large Language Models in Framing Detection
by: Pastorino, Valeria, et al.
Published: (2024)
by: Pastorino, Valeria, et al.
Published: (2024)
Hidden Failures in Robustness: Why Supervised Uncertainty Quantification Needs Better Evaluation
by: Stacey, Joe, et al.
Published: (2026)
by: Stacey, Joe, et al.
Published: (2026)
No Shortcuts to Culture: Indonesian Multi-hop Question Answering for Complex Cultural Understanding
by: Permadi, Vynska Amalia, et al.
Published: (2026)
by: Permadi, Vynska Amalia, et al.
Published: (2026)
Faithful Summarisation under Disagreement via Belief-Level Aggregation
by: Aghaebe, Favour Yahdii, et al.
Published: (2026)
by: Aghaebe, Favour Yahdii, et al.
Published: (2026)
Beyond Hate Speech: NLP's Challenges and Opportunities in Uncovering Dehumanizing Language
by: Saffari, Hamidreza, et al.
Published: (2024)
by: Saffari, Hamidreza, et al.
Published: (2024)
Exploring Gender Disparities in Automatic Speech Recognition Technology
by: ElGhazaly, Hend, et al.
Published: (2025)
by: ElGhazaly, Hend, et al.
Published: (2025)
Beyond One-Size-Fits-All: Inversion Learning for Highly Effective NLG Evaluation Prompts
by: Hong, Hanhua, et al.
Published: (2025)
by: Hong, Hanhua, et al.
Published: (2025)
Transforming Science with Large Language Models: A Survey on AI-assisted Scientific Discovery, Experimentation, Content Generation, and Evaluation
by: Eger, Steffen, et al.
Published: (2025)
by: Eger, Steffen, et al.
Published: (2025)
Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations
by: Roytburg, Dani, et al.
Published: (2026)
by: Roytburg, Dani, et al.
Published: (2026)
ReproHum #0087-01: Human Evaluation Reproduction Report for Generating Fact Checking Explanations
by: Loakman, Tyler, et al.
Published: (2024)
by: Loakman, Tyler, et al.
Published: (2024)
Emphasising Structured Information: Integrating Abstract Meaning Representation into LLMs for Enhanced Open-Domain Dialogue Evaluation
by: Yang, Bohao, et al.
Published: (2024)
by: Yang, Bohao, et al.
Published: (2024)
I Am Aligned, But With Whom? MENA Values Benchmark for Evaluating Cultural Alignment and Multilingual Bias in LLMs
by: Zahraei, Pardis Sadat, et al.
Published: (2025)
by: Zahraei, Pardis Sadat, et al.
Published: (2025)
Ara-HOPE: Human-Centric Post-Editing Evaluation for Dialectal Arabic to Modern Standard Arabic Translation
by: Alabdullah, Abdullah, et al.
Published: (2025)
by: Alabdullah, Abdullah, et al.
Published: (2025)
LatestEval: Addressing Data Contamination in Language Model Evaluation through Dynamic and Time-Sensitive Test Construction
by: Li, Yucheng, et al.
Published: (2023)
by: Li, Yucheng, et al.
Published: (2023)
Mirage of Mastery: Memorization Tricks LLMs into Artificially Inflated Self-Knowledge
by: Kale, Sahil
Published: (2025)
by: Kale, Sahil
Published: (2025)
Wired for Overconfidence: A Mechanistic Perspective on Inflated Verbalized Confidence in LLMs
by: Zhao, Tianyi, et al.
Published: (2026)
by: Zhao, Tianyi, et al.
Published: (2026)
From Facts to Insights: A Study on the Generation and Evaluation of Analytical Reports for Deciphering Earnings Calls
by: Goldsack, Tomas, et al.
Published: (2024)
by: Goldsack, Tomas, et al.
Published: (2024)
Evaluating Large Language Models for Generalization and Robustness via Data Compression
by: Li, Yucheng, et al.
Published: (2024)
by: Li, Yucheng, et al.
Published: (2024)
Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic Evaluation
by: Zhang, Xiangxu, et al.
Published: (2025)
by: Zhang, Xiangxu, et al.
Published: (2025)
PerCul: A Story-Driven Cultural Evaluation of LLMs in Persian
by: Monazzah, Erfan Moosavi, et al.
Published: (2025)
by: Monazzah, Erfan Moosavi, et al.
Published: (2025)
SLIDE: A Framework Integrating Small and Large Language Models for Open-Domain Dialogues Evaluation
by: Zhao, Kun, et al.
Published: (2024)
by: Zhao, Kun, et al.
Published: (2024)
MMTE: Corpus and Metrics for Evaluating Machine Translation Quality of Metaphorical Language
by: Wang, Shun, et al.
Published: (2024)
by: Wang, Shun, et al.
Published: (2024)
Effective Distillation of Table-based Reasoning Ability from LLMs
by: Yang, Bohao, et al.
Published: (2023)
by: Yang, Bohao, et al.
Published: (2023)
EgoDyn-Bench: Evaluating Ego-Motion Understanding in Vision-Centric Foundation Models for Autonomous Driving
by: Schäfer, Finn Rasmus, et al.
Published: (2026)
by: Schäfer, Finn Rasmus, et al.
Published: (2026)
When LLMs Struggle: Reference-less Translation Evaluation for Low-resource Languages
by: Sindhujan, Archchana, et al.
Published: (2025)
by: Sindhujan, Archchana, et al.
Published: (2025)
Encyclo-K: Evaluating LLMs with Dynamically Composed Knowledge Statements
by: Liang, Yiming, et al.
Published: (2025)
by: Liang, Yiming, et al.
Published: (2025)
SemScore: Automated Evaluation of Instruction-Tuned LLMs based on Semantic Textual Similarity
by: Aynetdinov, Ansar, et al.
Published: (2024)
by: Aynetdinov, Ansar, et al.
Published: (2024)
EgoThink: Evaluating First-Person Perspective Thinking Capability of Vision-Language Models
by: Cheng, Sijie, et al.
Published: (2023)
by: Cheng, Sijie, et al.
Published: (2023)
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
by: Xu, Wenda, et al.
Published: (2025)
by: Xu, Wenda, et al.
Published: (2025)
Similar Items
-
ContrastScore: Towards Higher Quality, Less Biased, More Efficient Evaluation Metrics with Contrastive Evaluation
by: Wang, Xiao, et al.
Published: (2025) -
RIGOURATE: Quantifying Scientific Exaggeration with Evidence-Aligned Claim Evaluation
by: James, Joseph, et al.
Published: (2026) -
Rolling the DICE on Idiomaticity: How LLMs Fail to Grasp Context
by: Mi, Maggie, et al.
Published: (2024) -
How to Leverage Digit Embeddings to Represent Numbers?
by: Sivakumar, Jasivan Alex, et al.
Published: (2024) -
Initialisation Determines the Basin: Efficient Codebook Optimisation for Extreme LLM Quantization
by: Kennedy, Ian W., et al.
Published: (2026)