Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | Molfese, Francesco Maria, Moroni, Luca, Gioffré, Luca, Scirè, Alessandro, Conia, Simone, Navigli, Roberto |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question Answering
by: Molfese, Francesco Maria, et al.
Published: (2025)
by: Molfese, Francesco Maria, et al.
Published: (2025)
ZEBRA: Zero-Shot Example-Based Retrieval Augmentation for Commonsense Question Answering
by: Molfese, Francesco Maria, et al.
Published: (2024)
by: Molfese, Francesco Maria, et al.
Published: (2024)
LiteraryQA: Towards Effective Evaluation of Long-document Narrative QA
by: Bonomo, Tommaso, et al.
Published: (2025)
by: Bonomo, Tommaso, et al.
Published: (2025)
FENICE: Factuality Evaluation of summarization based on Natural language Inference and Claim Extraction
by: Scirè, Alessandro, et al.
Published: (2024)
by: Scirè, Alessandro, et al.
Published: (2024)
Truth or Mirage? Towards End-to-End Factuality Evaluation with LLM-Oasis
by: Scirè, Alessandro, et al.
Published: (2024)
by: Scirè, Alessandro, et al.
Published: (2024)
Guardians of the Machine Translation Meta-Evaluation: Sentinel Metrics Fall In!
by: Perrella, Stefano, et al.
Published: (2024)
by: Perrella, Stefano, et al.
Published: (2024)
Improving Score Reliability of Multiple Choice Benchmarks with Consistency Evaluation and Altered Answer Choices
by: Cavalin, Paulo, et al.
Published: (2025)
by: Cavalin, Paulo, et al.
Published: (2025)
Polyglots or Multitudes? Multilingual LLM Answers to Value-laden Multiple-Choice Questions
by: Labat, Léo, et al.
Published: (2026)
by: Labat, Léo, et al.
Published: (2026)
LLM Distillation for Efficient Few-Shot Multiple Choice Question Answering
by: Sutanto, Patrick, et al.
Published: (2024)
by: Sutanto, Patrick, et al.
Published: (2024)
Differentiating Choices via Commonality for Multiple-Choice Question Answering
by: Deng, Wenqing, et al.
Published: (2024)
by: Deng, Wenqing, et al.
Published: (2024)
Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
by: Wiegreffe, Sarah, et al.
Published: (2024)
by: Wiegreffe, Sarah, et al.
Published: (2024)
Evaluating and Calibrating LLM Confidence on Questions with Multiple Correct Answers
by: Wang, Yuhan, et al.
Published: (2026)
by: Wang, Yuhan, et al.
Published: (2026)
Multiple-Choice Questions are Efficient and Robust LLM Evaluators
by: Zhang, Ziyin, et al.
Published: (2024)
by: Zhang, Ziyin, et al.
Published: (2024)
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
by: Balepur, Nishant, et al.
Published: (2024)
by: Balepur, Nishant, et al.
Published: (2024)
Improving LLM First-Token Predictions in Multiple-Choice Question Answering via Output Prefilling
by: Cappelletti, Silvia, et al.
Published: (2025)
by: Cappelletti, Silvia, et al.
Published: (2025)
Biomedical Entity Linking as Multiple Choice Question Answering
by: Lin, Zhenxi, et al.
Published: (2024)
by: Lin, Zhenxi, et al.
Published: (2024)
A Dataset of Open-Domain Question Answering with Multiple-Span Answers
by: Luo, Zhiyi, et al.
Published: (2024)
by: Luo, Zhiyi, et al.
Published: (2024)
LLMs Lost in Translation: M-ALERT uncovers Cross-Linguistic Safety Inconsistencies
by: Friedrich, Felix, et al.
Published: (2024)
by: Friedrich, Felix, et al.
Published: (2024)
Evaluating Answer Reranking Strategies in Time-sensitive Question Answering
by: Kardan, Mehmet, et al.
Published: (2025)
by: Kardan, Mehmet, et al.
Published: (2025)
Ensemble Transformer for Efficient and Accurate Ranking Tasks: an Application to Question Answering Systems
by: Matsubara, Yoshitomo, et al.
Published: (2022)
by: Matsubara, Yoshitomo, et al.
Published: (2022)
Consensus or Conflict? Fine-Grained Evaluation of Conflicting Answers in Question-Answering
by: Nachshoni, Eviatar, et al.
Published: (2025)
by: Nachshoni, Eviatar, et al.
Published: (2025)
Question Answering with LLMs and Learning from Answer Sets
by: Borroto, Manuel, et al.
Published: (2025)
by: Borroto, Manuel, et al.
Published: (2025)
Correctness Coverage Evaluation for Medical Multiple-Choice Question Answering Based on the Enhanced Conformal Prediction Framework
by: Ke, Yusong, et al.
Published: (2025)
by: Ke, Yusong, et al.
Published: (2025)
Applying Relation Extraction and Graph Matching to Answering Multiple Choice Questions
by: Shimoda, Naoki, et al.
Published: (2025)
by: Shimoda, Naoki, et al.
Published: (2025)
Anchored Answers: Unravelling Positional Bias in GPT-2's Multiple-Choice Questions
by: Li, Ruizhe, et al.
Published: (2024)
by: Li, Ruizhe, et al.
Published: (2024)
Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can't Answer?
by: Balepur, Nishant, et al.
Published: (2024)
by: Balepur, Nishant, et al.
Published: (2024)
Hierarchical Vision-Language Reasoning for Multimodal Multiple-Choice Question Answering
by: Zhou, Ao, et al.
Published: (2025)
by: Zhou, Ao, et al.
Published: (2025)
Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong
by: Fu, Tairan, et al.
Published: (2025)
by: Fu, Tairan, et al.
Published: (2025)
More Bias, Less Bias: BiasPrompting for Enhanced Multiple-Choice Question Answering
by: Vu, Duc Anh, et al.
Published: (2025)
by: Vu, Duc Anh, et al.
Published: (2025)
Mind the Gap: A Closer Look at Tokenization for Multiple-Choice Question Answering with LLMs
by: Sanz-Guerrero, Mario, et al.
Published: (2025)
by: Sanz-Guerrero, Mario, et al.
Published: (2025)
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
by: Chandak, Nikhil, et al.
Published: (2025)
by: Chandak, Nikhil, et al.
Published: (2025)
Meronymic Ontology Extraction via Large Language Models
by: Zhang, Dekai, et al.
Published: (2025)
by: Zhang, Dekai, et al.
Published: (2025)
Do Large Language Models Plan Answer Positions? Position Bias in Multiple-Choice Question Generation
by: Tang, Xuemei, et al.
Published: (2026)
by: Tang, Xuemei, et al.
Published: (2026)
Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
by: Yona, Gal, et al.
Published: (2024)
by: Yona, Gal, et al.
Published: (2024)
Varying Shades of Wrong: Aligning LLMs with Wrong Answers Only
by: Yao, Jihan, et al.
Published: (2024)
by: Yao, Jihan, et al.
Published: (2024)
A Study on Large Language Models' Limitations in Multiple-Choice Question Answering
by: Khatun, Aisha, et al.
Published: (2024)
by: Khatun, Aisha, et al.
Published: (2024)
Wrong Answers Can Also Be Useful: PlausibleQA -- A Large-Scale QA Dataset with Answer Plausibility Scores
by: Mozafari, Jamshid, et al.
Published: (2025)
by: Mozafari, Jamshid, et al.
Published: (2025)
Process Reward Models Meet Planning: Generating Precise and Scalable Datasets for Step-Level Rewards
by: Pisano, Raffaele, et al.
Published: (2026)
by: Pisano, Raffaele, et al.
Published: (2026)
Conformal P-Value in Multiple-Choice Question Answering Tasks with Provable Risk Control
by: Ye, Yuanchang
Published: (2025)
by: Ye, Yuanchang
Published: (2025)
Interpretable Coreference Resolution Evaluation Using Explicit Semantics
by: Gatti, Bruno, et al.
Published: (2026)
by: Gatti, Bruno, et al.
Published: (2026)
Similar Items
-
ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question Answering
by: Molfese, Francesco Maria, et al.
Published: (2025) -
ZEBRA: Zero-Shot Example-Based Retrieval Augmentation for Commonsense Question Answering
by: Molfese, Francesco Maria, et al.
Published: (2024) -
LiteraryQA: Towards Effective Evaluation of Long-document Narrative QA
by: Bonomo, Tommaso, et al.
Published: (2025) -
FENICE: Factuality Evaluation of summarization based on Natural language Inference and Claim Extraction
by: Scirè, Alessandro, et al.
Published: (2024) -
Truth or Mirage? Towards End-to-End Factuality Evaluation with LLM-Oasis
by: Scirè, Alessandro, et al.
Published: (2024)