Benchmarking is Broken -- Don't Let AI be its Own Judge
Fuente:
arXiv
Saved in:
| Main Authors: | Cheng, Zerui, Wohnig, Stella, Gupta, Ruchika, Alam, Samiul, Abdullahi, Tassallah, Ribeiro, João Alves, Nielsen-Garcia, Christian, Mir, Saif, Li, Siran, Orender, Jason, Bahrainian, Seyed Ali, Kirste, Daniel, Gokaslan, Aaron, Glinka, Mikołaj, Eickhoff, Carsten, Wolff, Ruben |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Retrieval Augmented Zero-Shot Text Classification
by: Abdullahi, Tassallah, et al.
Published: (2024)
by: Abdullahi, Tassallah, et al.
Published: (2024)
Enhancing Retrieval-Augmented Generation: A Study of Best Practices
by: Li, Siran, et al.
Published: (2025)
by: Li, Siran, et al.
Published: (2025)
Beyond Multiple Choice: Evaluating Steering Vectors for Summarization
by: Braun, Joschka, et al.
Published: (2025)
by: Braun, Joschka, et al.
Published: (2025)
When Silence Is Golden: Can LLMs Learn to Abstain in Temporal QA and Beyond?
by: Zhou, Xinyu, et al.
Published: (2026)
by: Zhou, Xinyu, et al.
Published: (2026)
Understanding (Un)Reliability of Steering Vectors in Language Models
by: Braun, Joschka, et al.
Published: (2025)
by: Braun, Joschka, et al.
Published: (2025)
UbuntuGuard: A Culturally-Grounded Policy Benchmark for Equitable AI Safety in African Languages
by: Abdullahi, Tassallah, et al.
Published: (2026)
by: Abdullahi, Tassallah, et al.
Published: (2026)
K-Paths: Reasoning over Graph Paths for Drug Repurposing and Drug Interaction Prediction
by: Abdullahi, Tassallah, et al.
Published: (2025)
by: Abdullahi, Tassallah, et al.
Published: (2025)
Don't Let MEV Slip: The Costs of Swapping on the Uniswap Protocol
by: Adams, Austin, et al.
Published: (2023)
by: Adams, Austin, et al.
Published: (2023)
The Persona Paradox: Medical Personas as Behavioral Priors in Clinical Language Models
by: Abdullahi, Tassallah, et al.
Published: (2026)
by: Abdullahi, Tassallah, et al.
Published: (2026)
When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't
by: Nemitz, Jonathan, et al.
Published: (2026)
by: Nemitz, Jonathan, et al.
Published: (2026)
Don't Judge by the Look: Towards Motion Coherent Video Representation
by: Zhang, Yitian, et al.
Published: (2024)
by: Zhang, Yitian, et al.
Published: (2024)
Don't Let a Few Network Failures Slow the Entire AllReduce
by: Chen, Peiqing, et al.
Published: (2026)
by: Chen, Peiqing, et al.
Published: (2026)
Don't Let It Hallucinate: Premise Verification via Retrieval-Augmented Logical Reasoning
by: Qin, Yuehan, et al.
Published: (2025)
by: Qin, Yuehan, et al.
Published: (2025)
F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare
by: Plyusov, Daniil, et al.
Published: (2026)
by: Plyusov, Daniil, et al.
Published: (2026)
Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
by: Moon, Jiwon, et al.
Published: (2025)
by: Moon, Jiwon, et al.
Published: (2025)
Logit Reweighting for Topic-Focused Summarization
by: Braun, Joschka, et al.
Published: (2025)
by: Braun, Joschka, et al.
Published: (2025)
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
by: Mayne, Harry, et al.
Published: (2025)
by: Mayne, Harry, et al.
Published: (2025)
Angles Don't Lie: Unlocking Training-Efficient RL Through the Model's Own Signals
by: Wang, Qinsi, et al.
Published: (2025)
by: Wang, Qinsi, et al.
Published: (2025)
Reach for Reference. Don't Judge a Database by Its Search Screen
by: Safford, Barbara Ripp
Published: (2005)
by: Safford, Barbara Ripp
Published: (2005)
Don't Let Bandit Feedback Pull Continual LLM-Recommender Updates Off Target
by: Kim, Taesan, et al.
Published: (2026)
by: Kim, Taesan, et al.
Published: (2026)
Don't Let Your Robot be Harmful: Responsible Robotic Manipulation via Safety-as-Policy
by: Ni, Minheng, et al.
Published: (2024)
by: Ni, Minheng, et al.
Published: (2024)
Don't Let Micropayments Penalize You--Experience from the City University of Hong Kong
by: Ching, Steve H., et al.
Published: (2009)
by: Ching, Steve H., et al.
Published: (2009)
FSCsec: Collaboration in Financial Sector Cybersecurity -- Exploring the Impact of Resource Sharing on IT Security
by: Sayeed, Sayed Abu, et al.
Published: (2024)
by: Sayeed, Sayed Abu, et al.
Published: (2024)
A Self-Attention-Driven Deep Denoiser Model for Real Time Lung Sound Denoising in Noisy Environments
by: Shuvo, Samiul Based, et al.
Published: (2024)
by: Shuvo, Samiul Based, et al.
Published: (2024)
Multiphoton-pumped UV-Vis transient absorption spectroscopy of 2D materials: basic concepts and recent applications
by: Glinka, Yuri D
Published: (2023)
by: Glinka, Yuri D
Published: (2023)
Ultrafast transient absorption spectroscopy of 2D semiconductors: a review
by: Glinka, Yuri D.
Published: (2025)
by: Glinka, Yuri D.
Published: (2025)
Let the Kids Choose Their Own Books.
by: Osgood, Thelma S.
Published: (1979)
by: Osgood, Thelma S.
Published: (1979)
Don't Let It Fade: Preserving Edits in Diffusion Language Models via Token Timestep Allocation
by: Kim, Woojin, et al.
Published: (2025)
by: Kim, Woojin, et al.
Published: (2025)
Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models
by: Baid, Ami, et al.
Published: (2026)
by: Baid, Ami, et al.
Published: (2026)
Don't Let Your COWS Be an Udder (Utter) Disaster: Wirelessness and Education Test the Electromagnetic Spectrum
by: Huber, Joe, et al.
Published: (2004)
by: Huber, Joe, et al.
Published: (2004)
Don't Judge Before You CLIP: A Unified Approach for Perceptual Tasks
by: Zalcher, Amit, et al.
Published: (2025)
by: Zalcher, Amit, et al.
Published: (2025)
Calibrate, Don't Curate: Label-Efficient Estimation from Noisy LLM Judges
by: Li, Yanran
Published: (2026)
by: Li, Yanran
Published: (2026)
Don't be salesmen
Published: (1997)
Published: (1997)
Don't Let the Claw Grip Your Hand: A Security Analysis and Defense Framework for OpenClaw
by: Shan, Zhengyang, et al.
Published: (2026)
by: Shan, Zhengyang, et al.
Published: (2026)
Navigating through the hidden embedding space: steering LLMs to improve mental health assessment
by: Ravenda, Federico, et al.
Published: (2025)
by: Ravenda, Federico, et al.
Published: (2025)
Trust, or Don't Predict: Introducing the CWSA Family for Confidence-Aware Model Evaluation
by: Shahnazari, Kourosh, et al.
Published: (2025)
by: Shahnazari, Kourosh, et al.
Published: (2025)
Don't Let AI Agents YOLO Your Files: Shifting Information and Control to Filesystems for Agent Safety and Autonomy
by: Zhong, Shawn Wanxiang, et al.
Published: (2026)
by: Zhong, Shawn Wanxiang, et al.
Published: (2026)
Don't Judge a Book by its Cover: Testing LLMs' Robustness Under Logical Obfuscation
by: Borah, Abhilekh, et al.
Published: (2026)
by: Borah, Abhilekh, et al.
Published: (2026)
Don’t Burn it Here
by: Walsh, Ed, et al.
Published: (2026)
by: Walsh, Ed, et al.
Published: (2026)
Don't Look Away
by: Cohen, Brianne
Published: (2023)
by: Cohen, Brianne
Published: (2023)
Similar Items
-
Retrieval Augmented Zero-Shot Text Classification
by: Abdullahi, Tassallah, et al.
Published: (2024) -
Enhancing Retrieval-Augmented Generation: A Study of Best Practices
by: Li, Siran, et al.
Published: (2025) -
Beyond Multiple Choice: Evaluating Steering Vectors for Summarization
by: Braun, Joschka, et al.
Published: (2025) -
When Silence Is Golden: Can LLMs Learn to Abstain in Temporal QA and Beyond?
by: Zhou, Xinyu, et al.
Published: (2026) -
Understanding (Un)Reliability of Steering Vectors in Language Models
by: Braun, Joschka, et al.
Published: (2025)