Gespeichert in:
| Hauptverfasser: | Salido, Eva Sánchez, Gonzalo, Julio, Marco, Guillermo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2502.12896 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Bilingual Evaluation of Language Models on General Knowledge in University Entrance Exams with Minimal Contamination
von: Salido, Eva Sánchez, et al.
Veröffentlicht: (2024)
von: Salido, Eva Sánchez, et al.
Veröffentlicht: (2024)
The Reader is the Metric: How Textual Features and Reader Profiles Explain Conflicting Evaluations of AI Creative Writing
von: Marco, Guillermo, et al.
Veröffentlicht: (2025)
von: Marco, Guillermo, et al.
Veröffentlicht: (2025)
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation
von: Hidayat, Naila Shafirni, et al.
Veröffentlicht: (2025)
von: Hidayat, Naila Shafirni, et al.
Veröffentlicht: (2025)
Multiple-Choice Questions are Efficient and Robust LLM Evaluators
von: Zhang, Ziyin, et al.
Veröffentlicht: (2024)
von: Zhang, Ziyin, et al.
Veröffentlicht: (2024)
Beyond Memorization: Distinguishing between Reductive and Epistemic Reasoning in LLMs using Classic Logic Puzzles
von: Gabay, Adi, et al.
Veröffentlicht: (2026)
von: Gabay, Adi, et al.
Veröffentlicht: (2026)
Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
von: Palta, Shramay, et al.
Veröffentlicht: (2024)
von: Palta, Shramay, et al.
Veröffentlicht: (2024)
None of the Above, Less of the Right: Parallel Patterns between Humans and LLMs on Multi-Choice Questions Answering
von: Tam, Zhi Rui, et al.
Veröffentlicht: (2025)
von: Tam, Zhi Rui, et al.
Veröffentlicht: (2025)
Small Language Models can Outperform Humans in Short Creative Writing: A Study Comparing SLMs with Humans and LLMs
von: Marco, Guillermo, et al.
Veröffentlicht: (2024)
von: Marco, Guillermo, et al.
Veröffentlicht: (2024)
Improving Score Reliability of Multiple Choice Benchmarks with Consistency Evaluation and Altered Answer Choices
von: Cavalin, Paulo, et al.
Veröffentlicht: (2025)
von: Cavalin, Paulo, et al.
Veröffentlicht: (2025)
Can LLM Graph Reasoning Generalize beyond Pattern Memorization?
von: Zhang, Yizhuo, et al.
Veröffentlicht: (2024)
von: Zhang, Yizhuo, et al.
Veröffentlicht: (2024)
BOE-XSUM: Extreme Summarization in Clear Language of Spanish Legal Decrees and Notifications
von: García, Andrés Fernández, et al.
Veröffentlicht: (2025)
von: García, Andrés Fernández, et al.
Veröffentlicht: (2025)
One Size Fits None: Heuristic Collapse in LLM Investment Advice
von: Ross, Jillian, et al.
Veröffentlicht: (2026)
von: Ross, Jillian, et al.
Veröffentlicht: (2026)
Distractor Generation in Multiple-Choice Tasks: A Survey of Methods, Datasets, and Evaluation
von: Alhazmi, Elaf, et al.
Veröffentlicht: (2024)
von: Alhazmi, Elaf, et al.
Veröffentlicht: (2024)
Benchmarking Chinese Commonsense Reasoning of LLMs: From Chinese-Specifics to Reasoning-Memorization Correlations
von: Sun, Jiaxing, et al.
Veröffentlicht: (2024)
von: Sun, Jiaxing, et al.
Veröffentlicht: (2024)
GeoChallenge: A Multi-Answer Multiple-Choice Benchmark for Geometric Reasoning with Diagrams
von: Zhang, Yushun, et al.
Veröffentlicht: (2026)
von: Zhang, Yushun, et al.
Veröffentlicht: (2026)
Reasoning Models are Test Exploiters: Rethinking Multiple-Choice
von: Raman, Narun, et al.
Veröffentlicht: (2025)
von: Raman, Narun, et al.
Veröffentlicht: (2025)
Enhancing Clinical Multiple-Choice Questions Benchmarks with Knowledge Graph Guided Distractor Generation
von: Yang, Running, et al.
Veröffentlicht: (2025)
von: Yang, Running, et al.
Veröffentlicht: (2025)
Beyond Memorization: Testing LLM Reasoning on Unseen Theory of Computation Tasks
von: Shelat, Shlok, et al.
Veröffentlicht: (2026)
von: Shelat, Shlok, et al.
Veröffentlicht: (2026)
Alleviating Choice Supportive Bias in LLM with Reasoning Dependency Generation
von: Zhuang, Nan, et al.
Veröffentlicht: (2025)
von: Zhuang, Nan, et al.
Veröffentlicht: (2025)
Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong
von: Fu, Tairan, et al.
Veröffentlicht: (2025)
von: Fu, Tairan, et al.
Veröffentlicht: (2025)
Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering
von: Molfese, Francesco Maria, et al.
Veröffentlicht: (2025)
von: Molfese, Francesco Maria, et al.
Veröffentlicht: (2025)
Reason to Rote: Rethinking Memorization in Reasoning
von: Du, Yupei, et al.
Veröffentlicht: (2025)
von: Du, Yupei, et al.
Veröffentlicht: (2025)
Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers
von: Balepur, Nishant, et al.
Veröffentlicht: (2025)
von: Balepur, Nishant, et al.
Veröffentlicht: (2025)
Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMs
von: Hans, Abhimanyu, et al.
Veröffentlicht: (2024)
von: Hans, Abhimanyu, et al.
Veröffentlicht: (2024)
Towards a Diagnostic and Predictive Evaluation Methodology for Sequence Labeling Tasks
von: Alvarez-Mellado, Elena, et al.
Veröffentlicht: (2026)
von: Alvarez-Mellado, Elena, et al.
Veröffentlicht: (2026)
Character-aware Transformers Learn an Irregular Morphological Pattern Yet None Generalize Like Humans
von: Ramarao, Akhilesh Kakolu, et al.
Veröffentlicht: (2026)
von: Ramarao, Akhilesh Kakolu, et al.
Veröffentlicht: (2026)
UBench: Benchmarking Uncertainty in Large Language Models with Multiple Choice Questions
von: Wang, Xunzhi, et al.
Veröffentlicht: (2024)
von: Wang, Xunzhi, et al.
Veröffentlicht: (2024)
Pron vs Prompt: Can Large Language Models already Challenge a World-Class Fiction Author at Creative Text Writing?
von: Marco, Guillermo, et al.
Veröffentlicht: (2024)
von: Marco, Guillermo, et al.
Veröffentlicht: (2024)
MatheMagic: Generating Dynamic Mathematics Benchmarks Robust to Memorization
von: O'Brien, Dayyán, et al.
Veröffentlicht: (2025)
von: O'Brien, Dayyán, et al.
Veröffentlicht: (2025)
AlpsBench: An LLM Personalization Benchmark for Real-Dialogue Memorization and Preference Alignment
von: Xiao, Jianfei, et al.
Veröffentlicht: (2026)
von: Xiao, Jianfei, et al.
Veröffentlicht: (2026)
Memorization or Interpolation ? Detecting LLM Memorization through Input Perturbation Analysis
von: Djiré, Albérick Euraste, et al.
Veröffentlicht: (2025)
von: Djiré, Albérick Euraste, et al.
Veröffentlicht: (2025)
Memorization or Reasoning? Exploring the Idiom Understanding of LLMs
von: Kim, Jisu, et al.
Veröffentlicht: (2025)
von: Kim, Jisu, et al.
Veröffentlicht: (2025)
On Memorization of Large Language Models in Logical Reasoning
von: Xie, Chulin, et al.
Veröffentlicht: (2024)
von: Xie, Chulin, et al.
Veröffentlicht: (2024)
LLM Distillation for Efficient Few-Shot Multiple Choice Question Answering
von: Sutanto, Patrick, et al.
Veröffentlicht: (2024)
von: Sutanto, Patrick, et al.
Veröffentlicht: (2024)
Pushing the Boundaries of Multiple Choice Evaluation to One Hundred Options
von: Lee, Nahyun, et al.
Veröffentlicht: (2026)
von: Lee, Nahyun, et al.
Veröffentlicht: (2026)
BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks
von: Balepur, Nishant, et al.
Veröffentlicht: (2026)
von: Balepur, Nishant, et al.
Veröffentlicht: (2026)
Beyond Multiple Choice: Evaluating Steering Vectors for Summarization
von: Braun, Joschka, et al.
Veröffentlicht: (2025)
von: Braun, Joschka, et al.
Veröffentlicht: (2025)
Polyglots or Multitudes? Multilingual LLM Answers to Value-laden Multiple-Choice Questions
von: Labat, Léo, et al.
Veröffentlicht: (2026)
von: Labat, Léo, et al.
Veröffentlicht: (2026)
Data Compressibility Quantifies LLM Memorization
von: Huang, Yizhan, et al.
Veröffentlicht: (2025)
von: Huang, Yizhan, et al.
Veröffentlicht: (2025)
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects
von: Mustapha, Ahmad, et al.
Veröffentlicht: (2024)
von: Mustapha, Ahmad, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Bilingual Evaluation of Language Models on General Knowledge in University Entrance Exams with Minimal Contamination
von: Salido, Eva Sánchez, et al.
Veröffentlicht: (2024) -
The Reader is the Metric: How Textual Features and Reader Profiles Explain Conflicting Evaluations of AI Creative Writing
von: Marco, Guillermo, et al.
Veröffentlicht: (2025) -
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation
von: Hidayat, Naila Shafirni, et al.
Veröffentlicht: (2025) -
Multiple-Choice Questions are Efficient and Robust LLM Evaluators
von: Zhang, Ziyin, et al.
Veröffentlicht: (2024) -
Beyond Memorization: Distinguishing between Reductive and Epistemic Reasoning in LLMs using Classic Logic Puzzles
von: Gabay, Adi, et al.
Veröffentlicht: (2026)