None of the Others: a General Technique to Distinguish Reasoning from Memorization in Multiple-Choice LLM Evaluation Benchmarks
Fuente:
arXiv
Guardado en:
| Autores principales: | Salido, Eva Sánchez, Gonzalo, Julio, Marco, Guillermo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation
por: Hidayat, Naila Shafirni, et al.
Publicado: (2025)
por: Hidayat, Naila Shafirni, et al.
Publicado: (2025)
Bilingual Evaluation of Language Models on General Knowledge in University Entrance Exams with Minimal Contamination
por: Salido, Eva Sánchez, et al.
Publicado: (2024)
por: Salido, Eva Sánchez, et al.
Publicado: (2024)
Multiple-Choice Questions are Efficient and Robust LLM Evaluators
por: Zhang, Ziyin, et al.
Publicado: (2024)
por: Zhang, Ziyin, et al.
Publicado: (2024)
Beyond Memorization: Distinguishing between Reductive and Epistemic Reasoning in LLMs using Classic Logic Puzzles
por: Gabay, Adi, et al.
Publicado: (2026)
por: Gabay, Adi, et al.
Publicado: (2026)
The Reader is the Metric: How Textual Features and Reader Profiles Explain Conflicting Evaluations of AI Creative Writing
por: Marco, Guillermo, et al.
Publicado: (2025)
por: Marco, Guillermo, et al.
Publicado: (2025)
Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
por: Palta, Shramay, et al.
Publicado: (2024)
por: Palta, Shramay, et al.
Publicado: (2024)
None of the Above, Less of the Right: Parallel Patterns between Humans and LLMs on Multi-Choice Questions Answering
por: Tam, Zhi Rui, et al.
Publicado: (2025)
por: Tam, Zhi Rui, et al.
Publicado: (2025)
Improving Score Reliability of Multiple Choice Benchmarks with Consistency Evaluation and Altered Answer Choices
por: Cavalin, Paulo, et al.
Publicado: (2025)
por: Cavalin, Paulo, et al.
Publicado: (2025)
Can LLM Graph Reasoning Generalize beyond Pattern Memorization?
por: Zhang, Yizhuo, et al.
Publicado: (2024)
por: Zhang, Yizhuo, et al.
Publicado: (2024)
Benchmarking Chinese Commonsense Reasoning of LLMs: From Chinese-Specifics to Reasoning-Memorization Correlations
por: Sun, Jiaxing, et al.
Publicado: (2024)
por: Sun, Jiaxing, et al.
Publicado: (2024)
Distractor Generation in Multiple-Choice Tasks: A Survey of Methods, Datasets, and Evaluation
por: Alhazmi, Elaf, et al.
Publicado: (2024)
por: Alhazmi, Elaf, et al.
Publicado: (2024)
One Size Fits None: Heuristic Collapse in LLM Investment Advice
por: Ross, Jillian, et al.
Publicado: (2026)
por: Ross, Jillian, et al.
Publicado: (2026)
GeoChallenge: A Multi-Answer Multiple-Choice Benchmark for Geometric Reasoning with Diagrams
por: Zhang, Yushun, et al.
Publicado: (2026)
por: Zhang, Yushun, et al.
Publicado: (2026)
Reasoning Models are Test Exploiters: Rethinking Multiple-Choice
por: Raman, Narun, et al.
Publicado: (2025)
por: Raman, Narun, et al.
Publicado: (2025)
Enhancing Clinical Multiple-Choice Questions Benchmarks with Knowledge Graph Guided Distractor Generation
por: Yang, Running, et al.
Publicado: (2025)
por: Yang, Running, et al.
Publicado: (2025)
Alleviating Choice Supportive Bias in LLM with Reasoning Dependency Generation
por: Zhuang, Nan, et al.
Publicado: (2025)
por: Zhuang, Nan, et al.
Publicado: (2025)
Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering
por: Molfese, Francesco Maria, et al.
Publicado: (2025)
por: Molfese, Francesco Maria, et al.
Publicado: (2025)
Beyond Memorization: Testing LLM Reasoning on Unseen Theory of Computation Tasks
por: Shelat, Shlok, et al.
Publicado: (2026)
por: Shelat, Shlok, et al.
Publicado: (2026)
Small Language Models can Outperform Humans in Short Creative Writing: A Study Comparing SLMs with Humans and LLMs
por: Marco, Guillermo, et al.
Publicado: (2024)
por: Marco, Guillermo, et al.
Publicado: (2024)
Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers
por: Balepur, Nishant, et al.
Publicado: (2025)
por: Balepur, Nishant, et al.
Publicado: (2025)
Reason to Rote: Rethinking Memorization in Reasoning
por: Du, Yupei, et al.
Publicado: (2025)
por: Du, Yupei, et al.
Publicado: (2025)
Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMs
por: Hans, Abhimanyu, et al.
Publicado: (2024)
por: Hans, Abhimanyu, et al.
Publicado: (2024)
UBench: Benchmarking Uncertainty in Large Language Models with Multiple Choice Questions
por: Wang, Xunzhi, et al.
Publicado: (2024)
por: Wang, Xunzhi, et al.
Publicado: (2024)
Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong
por: Fu, Tairan, et al.
Publicado: (2025)
por: Fu, Tairan, et al.
Publicado: (2025)
Character-aware Transformers Learn an Irregular Morphological Pattern Yet None Generalize Like Humans
por: Ramarao, Akhilesh Kakolu, et al.
Publicado: (2026)
por: Ramarao, Akhilesh Kakolu, et al.
Publicado: (2026)
MatheMagic: Generating Dynamic Mathematics Benchmarks Robust to Memorization
por: O'Brien, Dayyán, et al.
Publicado: (2025)
por: O'Brien, Dayyán, et al.
Publicado: (2025)
AlpsBench: An LLM Personalization Benchmark for Real-Dialogue Memorization and Preference Alignment
por: Xiao, Jianfei, et al.
Publicado: (2026)
por: Xiao, Jianfei, et al.
Publicado: (2026)
LLM Distillation for Efficient Few-Shot Multiple Choice Question Answering
por: Sutanto, Patrick, et al.
Publicado: (2024)
por: Sutanto, Patrick, et al.
Publicado: (2024)
Memorization or Reasoning? Exploring the Idiom Understanding of LLMs
por: Kim, Jisu, et al.
Publicado: (2025)
por: Kim, Jisu, et al.
Publicado: (2025)
On Memorization of Large Language Models in Logical Reasoning
por: Xie, Chulin, et al.
Publicado: (2024)
por: Xie, Chulin, et al.
Publicado: (2024)
Pushing the Boundaries of Multiple Choice Evaluation to One Hundred Options
por: Lee, Nahyun, et al.
Publicado: (2026)
por: Lee, Nahyun, et al.
Publicado: (2026)
BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks
por: Balepur, Nishant, et al.
Publicado: (2026)
por: Balepur, Nishant, et al.
Publicado: (2026)
Memorization or Interpolation ? Detecting LLM Memorization through Input Perturbation Analysis
por: Djiré, Albérick Euraste, et al.
Publicado: (2025)
por: Djiré, Albérick Euraste, et al.
Publicado: (2025)
BOE-XSUM: Extreme Summarization in Clear Language of Spanish Legal Decrees and Notifications
por: García, Andrés Fernández, et al.
Publicado: (2025)
por: García, Andrés Fernández, et al.
Publicado: (2025)
Polyglots or Multitudes? Multilingual LLM Answers to Value-laden Multiple-Choice Questions
por: Labat, Léo, et al.
Publicado: (2026)
por: Labat, Léo, et al.
Publicado: (2026)
Evaluating the Symbol Binding Ability of Large Language Models for Multiple-Choice Questions in Vietnamese General Education
por: Nguyen, Duc-Vu, et al.
Publicado: (2023)
por: Nguyen, Duc-Vu, et al.
Publicado: (2023)
Beyond Multiple Choice: Evaluating Steering Vectors for Summarization
por: Braun, Joschka, et al.
Publicado: (2025)
por: Braun, Joschka, et al.
Publicado: (2025)
Unveiling Over-Memorization in Finetuning LLMs for Reasoning Tasks
por: Ruan, Zhiwen, et al.
Publicado: (2025)
por: Ruan, Zhiwen, et al.
Publicado: (2025)
MCS-SQL: Leveraging Multiple Prompts and Multiple-Choice Selection For Text-to-SQL Generation
por: Lee, Dongjun, et al.
Publicado: (2024)
por: Lee, Dongjun, et al.
Publicado: (2024)
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects
por: Mustapha, Ahmad, et al.
Publicado: (2024)
por: Mustapha, Ahmad, et al.
Publicado: (2024)
Ejemplares similares
-
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation
por: Hidayat, Naila Shafirni, et al.
Publicado: (2025) -
Bilingual Evaluation of Language Models on General Knowledge in University Entrance Exams with Minimal Contamination
por: Salido, Eva Sánchez, et al.
Publicado: (2024) -
Multiple-Choice Questions are Efficient and Robust LLM Evaluators
por: Zhang, Ziyin, et al.
Publicado: (2024) -
Beyond Memorization: Distinguishing between Reductive and Epistemic Reasoning in LLMs using Classic Logic Puzzles
por: Gabay, Adi, et al.
Publicado: (2026) -
The Reader is the Metric: How Textual Features and Reader Profiles Explain Conflicting Evaluations of AI Creative Writing
por: Marco, Guillermo, et al.
Publicado: (2025)