The Base-Rate Effect on LLM Benchmark Performance: Disambiguating Test-Taking Strategies from Benchmark Performance
Fuente:
arXiv
Saved in:
| Main Authors: | Moore, Kyle, Roberts, Jesse, Pham, Thao, Ewaleifoh, Oseremhen, Fisher, Doug |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Large Language Model Recall Uncertainty is Modulated by the Fan Effect
by: Roberts, Jesse, et al.
Published: (2024)
by: Roberts, Jesse, et al.
Published: (2024)
Chain of Thought Still Thinks Fast: APriCoT Helps with Thinking Slow
by: Moore, Kyle, et al.
Published: (2024)
by: Moore, Kyle, et al.
Published: (2024)
Do Large Language Models Learn Human-Like Strategic Preferences?
by: Roberts, Jesse, et al.
Published: (2024)
by: Roberts, Jesse, et al.
Published: (2024)
Human-Alignment and Calibration of Inference-Time Uncertainty in Large Language Models
by: Moore, Kyle, et al.
Published: (2025)
by: Moore, Kyle, et al.
Published: (2025)
Scheming Ability in LLM-to-LLM Strategic Interactions
by: Pham, Thao
Published: (2025)
by: Pham, Thao
Published: (2025)
Using Artificial Populations to Study Psychological Phenomena in Neural Models
by: Roberts, Jesse, et al.
Published: (2023)
by: Roberts, Jesse, et al.
Published: (2023)
NC-Bench: An LLM Benchmark for Evaluating Conversational Competence
by: Moore, Robert J., et al.
Published: (2026)
by: Moore, Robert J., et al.
Published: (2026)
Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment
by: Atasoy, I. F., et al.
Published: (2026)
by: Atasoy, I. F., et al.
Published: (2026)
Benchmark Test-Time Scaling of General LLM Agents
by: Li, Xiaochuan, et al.
Published: (2026)
by: Li, Xiaochuan, et al.
Published: (2026)
Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts
by: Haimes, Jacob, et al.
Published: (2024)
by: Haimes, Jacob, et al.
Published: (2024)
Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing
by: Berghaus, David, et al.
Published: (2025)
by: Berghaus, David, et al.
Published: (2025)
Evaluating the Performance of Large Language Models on GAOKAO Benchmark
by: Zhang, Xiaotian, et al.
Published: (2023)
by: Zhang, Xiaotian, et al.
Published: (2023)
Benchmark Profiling: Mechanistic Diagnosis of LLM Benchmarks
by: Kim, Dongjun, et al.
Published: (2025)
by: Kim, Dongjun, et al.
Published: (2025)
Query Disambiguation via Answer-Free Context: Doubling Performance on Humanity's Last Exam
by: Majurski, Michael, et al.
Published: (2026)
by: Majurski, Michael, et al.
Published: (2026)
Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
by: Li, Peiyu, et al.
Published: (2025)
by: Li, Peiyu, et al.
Published: (2025)
From Test-Taking to Test-Making: Examining LLM Authoring of Commonsense Assessment Items
by: Roemmele, Melissa, et al.
Published: (2024)
by: Roemmele, Melissa, et al.
Published: (2024)
Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap
by: Srivastava, Saurabh, et al.
Published: (2024)
by: Srivastava, Saurabh, et al.
Published: (2024)
LLM CHESS: Benchmarking Reasoning and Instruction-Following in LLMs through Chess
by: Kolasani, Sai, et al.
Published: (2025)
by: Kolasani, Sai, et al.
Published: (2025)
Measuring Competency, Not Performance: Item-Aware Evaluation Across Medical Benchmarks
by: Luo, Zhimeng, et al.
Published: (2025)
by: Luo, Zhimeng, et al.
Published: (2025)
Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
by: Wei, Tianxin, et al.
Published: (2025)
by: Wei, Tianxin, et al.
Published: (2025)
Who Benchmarks the Benchmarks? A Case Study of LLM Evaluation in Icelandic
by: Ingimundarson, Finnur Ágúst, et al.
Published: (2026)
by: Ingimundarson, Finnur Ágúst, et al.
Published: (2026)
Teaching Language Models to Check Grounded Claim Factuality with Human Test-Taking Strategies
by: Ye, Yuxuan, et al.
Published: (2026)
by: Ye, Yuxuan, et al.
Published: (2026)
The Emperor's New Clothes in Benchmarking? A Rigorous Examination of Mitigation Strategies for LLM Benchmark Data Contamination
by: Sun, Yifan, et al.
Published: (2025)
by: Sun, Yifan, et al.
Published: (2025)
Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic Evaluation
by: Zhang, Xiangxu, et al.
Published: (2025)
by: Zhang, Xiangxu, et al.
Published: (2025)
Fluid Language Model Benchmarking
by: Hofmann, Valentin, et al.
Published: (2025)
by: Hofmann, Valentin, et al.
Published: (2025)
Silicon Bureaucracy and AI Test-Oriented Education: Contamination Sensitivity and Score Confidence in LLM Benchmarks
by: Song, Yiliang, et al.
Published: (2026)
by: Song, Yiliang, et al.
Published: (2026)
HalluLens: LLM Hallucination Benchmark
by: Bang, Yejin, et al.
Published: (2025)
by: Bang, Yejin, et al.
Published: (2025)
Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling
by: Sun, Shengyin, et al.
Published: (2025)
by: Sun, Shengyin, et al.
Published: (2025)
VN-MTEB: Vietnamese Massive Text Embedding Benchmark
by: Pham, Loc, et al.
Published: (2025)
by: Pham, Loc, et al.
Published: (2025)
Self-Reflection in LLM Agents: Effects on Problem-Solving Performance
by: Renze, Matthew, et al.
Published: (2024)
by: Renze, Matthew, et al.
Published: (2024)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
by: Tan, Sijun, et al.
Published: (2024)
by: Tan, Sijun, et al.
Published: (2024)
Language Models as Knowledge Bases for Visual Word Sense Disambiguation
by: Kritharoula, Anastasia, et al.
Published: (2023)
by: Kritharoula, Anastasia, et al.
Published: (2023)
Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards
by: Tamber, Manveer Singh, et al.
Published: (2025)
by: Tamber, Manveer Singh, et al.
Published: (2025)
Spanish and LLM Benchmarks: is MMLU Lost in Translation?
by: Plaza, Irene, et al.
Published: (2024)
by: Plaza, Irene, et al.
Published: (2024)
Benchmark of stylistic variation in LLM-generated texts
by: Milička, Jiří, et al.
Published: (2025)
by: Milička, Jiří, et al.
Published: (2025)
Benchmarking and Improving LLM Robustness for Personalized Generation
by: Okite, Chimaobi, et al.
Published: (2025)
by: Okite, Chimaobi, et al.
Published: (2025)
PreScience: A Benchmark for Forecasting Scientific Contributions
by: Ajith, Anirudh, et al.
Published: (2026)
by: Ajith, Anirudh, et al.
Published: (2026)
Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs
by: Restrepo, David, et al.
Published: (2024)
by: Restrepo, David, et al.
Published: (2024)
BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law
by: Nagl, Sebastian, et al.
Published: (2026)
by: Nagl, Sebastian, et al.
Published: (2026)
LLMs as Agentic Cooperative Players in Multiplayer UNO
by: Matinez, Yago Romano, et al.
Published: (2025)
by: Matinez, Yago Romano, et al.
Published: (2025)
Similar Items
-
Large Language Model Recall Uncertainty is Modulated by the Fan Effect
by: Roberts, Jesse, et al.
Published: (2024) -
Chain of Thought Still Thinks Fast: APriCoT Helps with Thinking Slow
by: Moore, Kyle, et al.
Published: (2024) -
Do Large Language Models Learn Human-Like Strategic Preferences?
by: Roberts, Jesse, et al.
Published: (2024) -
Human-Alignment and Calibration of Inference-Time Uncertainty in Large Language Models
by: Moore, Kyle, et al.
Published: (2025) -
Scheming Ability in LLM-to-LLM Strategic Interactions
by: Pham, Thao
Published: (2025)