Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Haimes, Jacob, Wenner, Cenny, Thaman, Kunvar, Tashev, Vassil, Neo, Clement, Kran, Esben, Schreiber, Jason |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use
von: Thaman, Kunvar
Veröffentlicht: (2026)
von: Thaman, Kunvar
Veröffentlicht: (2026)
Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities
von: Anurin, Andrey, et al.
Veröffentlicht: (2024)
von: Anurin, Andrey, et al.
Veröffentlicht: (2024)
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
von: Veuthey, Jaime Raldua, et al.
Veröffentlicht: (2025)
von: Veuthey, Jaime Raldua, et al.
Veröffentlicht: (2025)
DarkBench: Benchmarking Dark Patterns in Large Language Models
von: Kran, Esben, et al.
Veröffentlicht: (2025)
von: Kran, Esben, et al.
Veröffentlicht: (2025)
View From Above: A Framework for Evaluating Distribution Shifts in Model Behavior
von: Chopra, Tanush, et al.
Veröffentlicht: (2024)
von: Chopra, Tanush, et al.
Veröffentlicht: (2024)
Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis
von: Reese, May Lynn, et al.
Veröffentlicht: (2026)
von: Reese, May Lynn, et al.
Veröffentlicht: (2026)
Tailored Truths: Optimizing LLM Persuasion with Personalization and Fabricated Statistics
von: Timm, Jasper, et al.
Veröffentlicht: (2025)
von: Timm, Jasper, et al.
Veröffentlicht: (2025)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
von: Neo, Clement, et al.
Veröffentlicht: (2024)
von: Neo, Clement, et al.
Veröffentlicht: (2024)
HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
von: Sturgeon, Benjamin, et al.
Veröffentlicht: (2025)
von: Sturgeon, Benjamin, et al.
Veröffentlicht: (2025)
Approximating Human Preferences Using a Multi-Judge Learned System
von: Sprejer, Eitán, et al.
Veröffentlicht: (2025)
von: Sprejer, Eitán, et al.
Veröffentlicht: (2025)
Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models
von: Tice, Cameron, et al.
Veröffentlicht: (2024)
von: Tice, Cameron, et al.
Veröffentlicht: (2024)
Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap
von: Srivastava, Saurabh, et al.
Veröffentlicht: (2024)
von: Srivastava, Saurabh, et al.
Veröffentlicht: (2024)
RetroLLM: Empowering Large Language Models to Retrieve Fine-grained Evidence within Generation
von: Li, Xiaoxi, et al.
Veröffentlicht: (2024)
von: Li, Xiaoxi, et al.
Veröffentlicht: (2024)
DeepRetro: Retrosynthetic Pathway Discovery using Iterative LLM Reasoning
von: Sathyanarayana, Shreyas Vinaya, et al.
Veröffentlicht: (2025)
von: Sathyanarayana, Shreyas Vinaya, et al.
Veröffentlicht: (2025)
HealthQA-BR: A System-Wide Benchmark Reveals Critical Knowledge Gaps in Large Language Models
von: D'addario, Andrew Maranhão Ventura
Veröffentlicht: (2025)
von: D'addario, Andrew Maranhão Ventura
Veröffentlicht: (2025)
Retro*: Optimizing LLMs for Reasoning-Intensive Document Retrieval
von: Lan, Junwei, et al.
Veröffentlicht: (2025)
von: Lan, Junwei, et al.
Veröffentlicht: (2025)
Bridging the Gap: Enhancing LLM Performance for Low-Resource African Languages with New Benchmarks, Fine-Tuning, and Cultural Adjustments
von: Alhanai, Tuka, et al.
Veröffentlicht: (2024)
von: Alhanai, Tuka, et al.
Veröffentlicht: (2024)
The Base-Rate Effect on LLM Benchmark Performance: Disambiguating Test-Taking Strategies from Benchmark Performance
von: Moore, Kyle, et al.
Veröffentlicht: (2024)
von: Moore, Kyle, et al.
Veröffentlicht: (2024)
Retro-Search: Exploring Untaken Paths for Deeper and Efficient Reasoning
von: Lu, Ximing, et al.
Veröffentlicht: (2025)
von: Lu, Ximing, et al.
Veröffentlicht: (2025)
Rethinking CyberSecEval: An LLM-Aided Approach to Evaluation Critique
von: Hariharan, Suhas, et al.
Veröffentlicht: (2024)
von: Hariharan, Suhas, et al.
Veröffentlicht: (2024)
Phonetic Perturbations Reveal Tokenizer-Rooted Safety Gaps in LLMs
von: Aswal, Darpan, et al.
Veröffentlicht: (2025)
von: Aswal, Darpan, et al.
Veröffentlicht: (2025)
Revealing the Numeracy Gap: An Empirical Investigation of Text Embedding Models
von: Deng, Ningyuan, et al.
Veröffentlicht: (2025)
von: Deng, Ningyuan, et al.
Veröffentlicht: (2025)
Holdout-Loss-Based Data Selection for LLM Finetuning via In-Context Learning
von: Zhang, Ling, et al.
Veröffentlicht: (2025)
von: Zhang, Ling, et al.
Veröffentlicht: (2025)
Exposing Numeracy Gaps: A Benchmark to Evaluate Fundamental Numerical Abilities in Large Language Models
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
LLM Unlearning Reveals a Stronger-Than-Expected Coreset Effect in Current Benchmarks
von: Pal, Soumyadeep, et al.
Veröffentlicht: (2025)
von: Pal, Soumyadeep, et al.
Veröffentlicht: (2025)
InstructRetro: Instruction Tuning post Retrieval-Augmented Pretraining
von: Wang, Boxin, et al.
Veröffentlicht: (2023)
von: Wang, Boxin, et al.
Veröffentlicht: (2023)
Bridging Linguistic Gaps: Cross-Lingual Mapping in Pre-Training and Dataset for Enhanced Multilingual LLM Performance
von: Zheng, Weihua, et al.
Veröffentlicht: (2026)
von: Zheng, Weihua, et al.
Veröffentlicht: (2026)
Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment
von: Atasoy, I. F., et al.
Veröffentlicht: (2026)
von: Atasoy, I. F., et al.
Veröffentlicht: (2026)
Causal Interventions Reveal Shared Structure Across English Filler-Gap Constructions
von: Boguraev, Sasha, et al.
Veröffentlicht: (2025)
von: Boguraev, Sasha, et al.
Veröffentlicht: (2025)
Multi-Agent Security Tax: Trading Off Security and Collaboration Capabilities in Multi-Agent Systems
von: Peigne-Lefebvre, Pierre, et al.
Veröffentlicht: (2025)
von: Peigne-Lefebvre, Pierre, et al.
Veröffentlicht: (2025)
Answering Students' Questions on Course Forums Using Multiple Chain-of-Thought Reasoning and Finetuning RAG-Enabled LLM
von: Wang, Neo, et al.
Veröffentlicht: (2025)
von: Wang, Neo, et al.
Veröffentlicht: (2025)
User-LLM: Efficient LLM Contextualization with User Embeddings
von: Ning, Lin, et al.
Veröffentlicht: (2024)
von: Ning, Lin, et al.
Veröffentlicht: (2024)
Benchmark Profiling: Mechanistic Diagnosis of LLM Benchmarks
von: Kim, Dongjun, et al.
Veröffentlicht: (2025)
von: Kim, Dongjun, et al.
Veröffentlicht: (2025)
Separating Tongue from Thought: Activation Patching Reveals Language-Agnostic Concept Representations in Transformers
von: Dumas, Clément, et al.
Veröffentlicht: (2024)
von: Dumas, Clément, et al.
Veröffentlicht: (2024)
Understanding Addition and Subtraction in Transformers
von: Quirke, Philip, et al.
Veröffentlicht: (2024)
von: Quirke, Philip, et al.
Veröffentlicht: (2024)
UserSumBench: A Benchmark Framework for Evaluating User Summarization Approaches
von: Wang, Chao, et al.
Veröffentlicht: (2024)
von: Wang, Chao, et al.
Veröffentlicht: (2024)
BEExAI: Benchmark to Evaluate Explainable AI
von: Sithakoul, Samuel, et al.
Veröffentlicht: (2024)
von: Sithakoul, Samuel, et al.
Veröffentlicht: (2024)
Mind the Gap: The Divergence Between Human and LLM-Generated Tasks
von: Lu, Yi-Long, et al.
Veröffentlicht: (2025)
von: Lu, Yi-Long, et al.
Veröffentlicht: (2025)
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
von: Skorobogat, Ronald, et al.
Veröffentlicht: (2026)
von: Skorobogat, Ronald, et al.
Veröffentlicht: (2026)
Smooth Reading: Bridging the Gap of Recurrent LLM to Self-Attention LLM on Long-Context Tasks
von: Liu, Kai, et al.
Veröffentlicht: (2025)
von: Liu, Kai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use
von: Thaman, Kunvar
Veröffentlicht: (2026) -
Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities
von: Anurin, Andrey, et al.
Veröffentlicht: (2024) -
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
von: Veuthey, Jaime Raldua, et al.
Veröffentlicht: (2025) -
DarkBench: Benchmarking Dark Patterns in Large Language Models
von: Kran, Esben, et al.
Veröffentlicht: (2025) -
View From Above: A Framework for Evaluating Distribution Shifts in Model Behavior
von: Chopra, Tanush, et al.
Veröffentlicht: (2024)