The Base-Rate Effect on LLM Benchmark Performance: Disambiguating Test-Taking Strategies from Benchmark Performance
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Moore, Kyle, Roberts, Jesse, Pham, Thao, Ewaleifoh, Oseremhen, Fisher, Doug |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Large Language Model Recall Uncertainty is Modulated by the Fan Effect
par: Roberts, Jesse, et autres
Publié: (2024)
par: Roberts, Jesse, et autres
Publié: (2024)
Chain of Thought Still Thinks Fast: APriCoT Helps with Thinking Slow
par: Moore, Kyle, et autres
Publié: (2024)
par: Moore, Kyle, et autres
Publié: (2024)
Do Large Language Models Learn Human-Like Strategic Preferences?
par: Roberts, Jesse, et autres
Publié: (2024)
par: Roberts, Jesse, et autres
Publié: (2024)
Human-Alignment and Calibration of Inference-Time Uncertainty in Large Language Models
par: Moore, Kyle, et autres
Publié: (2025)
par: Moore, Kyle, et autres
Publié: (2025)
Scheming Ability in LLM-to-LLM Strategic Interactions
par: Pham, Thao
Publié: (2025)
par: Pham, Thao
Publié: (2025)
Using Artificial Populations to Study Psychological Phenomena in Neural Models
par: Roberts, Jesse, et autres
Publié: (2023)
par: Roberts, Jesse, et autres
Publié: (2023)
NC-Bench: An LLM Benchmark for Evaluating Conversational Competence
par: Moore, Robert J., et autres
Publié: (2026)
par: Moore, Robert J., et autres
Publié: (2026)
Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment
par: Atasoy, I. F., et autres
Publié: (2026)
par: Atasoy, I. F., et autres
Publié: (2026)
Benchmark Test-Time Scaling of General LLM Agents
par: Li, Xiaochuan, et autres
Publié: (2026)
par: Li, Xiaochuan, et autres
Publié: (2026)
Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts
par: Haimes, Jacob, et autres
Publié: (2024)
par: Haimes, Jacob, et autres
Publié: (2024)
Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing
par: Berghaus, David, et autres
Publié: (2025)
par: Berghaus, David, et autres
Publié: (2025)
Evaluating the Performance of Large Language Models on GAOKAO Benchmark
par: Zhang, Xiaotian, et autres
Publié: (2023)
par: Zhang, Xiaotian, et autres
Publié: (2023)
Benchmark Profiling: Mechanistic Diagnosis of LLM Benchmarks
par: Kim, Dongjun, et autres
Publié: (2025)
par: Kim, Dongjun, et autres
Publié: (2025)
Query Disambiguation via Answer-Free Context: Doubling Performance on Humanity's Last Exam
par: Majurski, Michael, et autres
Publié: (2026)
par: Majurski, Michael, et autres
Publié: (2026)
Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
par: Li, Peiyu, et autres
Publié: (2025)
par: Li, Peiyu, et autres
Publié: (2025)
From Test-Taking to Test-Making: Examining LLM Authoring of Commonsense Assessment Items
par: Roemmele, Melissa, et autres
Publié: (2024)
par: Roemmele, Melissa, et autres
Publié: (2024)
Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap
par: Srivastava, Saurabh, et autres
Publié: (2024)
par: Srivastava, Saurabh, et autres
Publié: (2024)
LLM CHESS: Benchmarking Reasoning and Instruction-Following in LLMs through Chess
par: Kolasani, Sai, et autres
Publié: (2025)
par: Kolasani, Sai, et autres
Publié: (2025)
Measuring Competency, Not Performance: Item-Aware Evaluation Across Medical Benchmarks
par: Luo, Zhimeng, et autres
Publié: (2025)
par: Luo, Zhimeng, et autres
Publié: (2025)
Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
par: Wei, Tianxin, et autres
Publié: (2025)
par: Wei, Tianxin, et autres
Publié: (2025)
Who Benchmarks the Benchmarks? A Case Study of LLM Evaluation in Icelandic
par: Ingimundarson, Finnur Ágúst, et autres
Publié: (2026)
par: Ingimundarson, Finnur Ágúst, et autres
Publié: (2026)
Teaching Language Models to Check Grounded Claim Factuality with Human Test-Taking Strategies
par: Ye, Yuxuan, et autres
Publié: (2026)
par: Ye, Yuxuan, et autres
Publié: (2026)
The Emperor's New Clothes in Benchmarking? A Rigorous Examination of Mitigation Strategies for LLM Benchmark Data Contamination
par: Sun, Yifan, et autres
Publié: (2025)
par: Sun, Yifan, et autres
Publié: (2025)
Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic Evaluation
par: Zhang, Xiangxu, et autres
Publié: (2025)
par: Zhang, Xiangxu, et autres
Publié: (2025)
Fluid Language Model Benchmarking
par: Hofmann, Valentin, et autres
Publié: (2025)
par: Hofmann, Valentin, et autres
Publié: (2025)
Silicon Bureaucracy and AI Test-Oriented Education: Contamination Sensitivity and Score Confidence in LLM Benchmarks
par: Song, Yiliang, et autres
Publié: (2026)
par: Song, Yiliang, et autres
Publié: (2026)
HalluLens: LLM Hallucination Benchmark
par: Bang, Yejin, et autres
Publié: (2025)
par: Bang, Yejin, et autres
Publié: (2025)
Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling
par: Sun, Shengyin, et autres
Publié: (2025)
par: Sun, Shengyin, et autres
Publié: (2025)
VN-MTEB: Vietnamese Massive Text Embedding Benchmark
par: Pham, Loc, et autres
Publié: (2025)
par: Pham, Loc, et autres
Publié: (2025)
Self-Reflection in LLM Agents: Effects on Problem-Solving Performance
par: Renze, Matthew, et autres
Publié: (2024)
par: Renze, Matthew, et autres
Publié: (2024)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
par: Tan, Sijun, et autres
Publié: (2024)
par: Tan, Sijun, et autres
Publié: (2024)
Language Models as Knowledge Bases for Visual Word Sense Disambiguation
par: Kritharoula, Anastasia, et autres
Publié: (2023)
par: Kritharoula, Anastasia, et autres
Publié: (2023)
Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards
par: Tamber, Manveer Singh, et autres
Publié: (2025)
par: Tamber, Manveer Singh, et autres
Publié: (2025)
Spanish and LLM Benchmarks: is MMLU Lost in Translation?
par: Plaza, Irene, et autres
Publié: (2024)
par: Plaza, Irene, et autres
Publié: (2024)
Benchmark of stylistic variation in LLM-generated texts
par: Milička, Jiří, et autres
Publié: (2025)
par: Milička, Jiří, et autres
Publié: (2025)
Benchmarking and Improving LLM Robustness for Personalized Generation
par: Okite, Chimaobi, et autres
Publié: (2025)
par: Okite, Chimaobi, et autres
Publié: (2025)
PreScience: A Benchmark for Forecasting Scientific Contributions
par: Ajith, Anirudh, et autres
Publié: (2026)
par: Ajith, Anirudh, et autres
Publié: (2026)
Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs
par: Restrepo, David, et autres
Publié: (2024)
par: Restrepo, David, et autres
Publié: (2024)
BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law
par: Nagl, Sebastian, et autres
Publié: (2026)
par: Nagl, Sebastian, et autres
Publié: (2026)
LLMs as Agentic Cooperative Players in Multiplayer UNO
par: Matinez, Yago Romano, et autres
Publié: (2025)
par: Matinez, Yago Romano, et autres
Publié: (2025)
Documents similaires
-
Large Language Model Recall Uncertainty is Modulated by the Fan Effect
par: Roberts, Jesse, et autres
Publié: (2024) -
Chain of Thought Still Thinks Fast: APriCoT Helps with Thinking Slow
par: Moore, Kyle, et autres
Publié: (2024) -
Do Large Language Models Learn Human-Like Strategic Preferences?
par: Roberts, Jesse, et autres
Publié: (2024) -
Human-Alignment and Calibration of Inference-Time Uncertainty in Large Language Models
par: Moore, Kyle, et autres
Publié: (2025) -
Scheming Ability in LLM-to-LLM Strategic Interactions
par: Pham, Thao
Publié: (2025)