BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks
Fuente:
arXiv
Saved in:
| Main Authors: | Balepur, Nishant, Rajasekaran, Bhavya, Oh, Jane, Xie, Michael, Desai, Atrey, Gupta, Vipul, Moore, Steven James, Choi, Eunsol, Rudinger, Rachel, Boyd-Graber, Jordan Lee |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
by: Balepur, Nishant, et al.
Published: (2025)
by: Balepur, Nishant, et al.
Published: (2025)
Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers
by: Balepur, Nishant, et al.
Published: (2025)
by: Balepur, Nishant, et al.
Published: (2025)
Is Your Large Language Model Knowledgeable or a Choices-Only Cheater?
by: Balepur, Nishant, et al.
Published: (2024)
by: Balepur, Nishant, et al.
Published: (2024)
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
by: Balepur, Nishant, et al.
Published: (2024)
by: Balepur, Nishant, et al.
Published: (2024)
Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can't Answer?
by: Balepur, Nishant, et al.
Published: (2024)
by: Balepur, Nishant, et al.
Published: (2024)
Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User Personas
by: Balepur, Nishant, et al.
Published: (2025)
by: Balepur, Nishant, et al.
Published: (2025)
It's Not Easy Being Wrong: Large Language Models Struggle with Process of Elimination Reasoning
by: Balepur, Nishant, et al.
Published: (2023)
by: Balepur, Nishant, et al.
Published: (2023)
KARL: Knowledge-Aware Retrieval and Representations aid Retention and Learning in Students
by: Shu, Matthew, et al.
Published: (2024)
by: Shu, Matthew, et al.
Published: (2024)
DiscoTrace: Representing and Comparing Answering Strategies of Humans and LLMs in Information-Seeking Question Answering
by: Srikanth, Neha, et al.
Published: (2026)
by: Srikanth, Neha, et al.
Published: (2026)
Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
by: Palta, Shramay, et al.
Published: (2024)
by: Palta, Shramay, et al.
Published: (2024)
A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
by: Balepur, Nishant, et al.
Published: (2025)
by: Balepur, Nishant, et al.
Published: (2025)
DRACULA: Hunting for the Actions Users Want Deep Research Agents to Execute
by: Balepur, Nishant, et al.
Published: (2026)
by: Balepur, Nishant, et al.
Published: (2026)
MODS: Moderating a Mixture of Document Speakers to Summarize Debatable Queries in Document Collections
by: Balepur, Nishant, et al.
Published: (2025)
by: Balepur, Nishant, et al.
Published: (2025)
Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users
by: Balepur, Nishant, et al.
Published: (2026)
by: Balepur, Nishant, et al.
Published: (2026)
A SMART Mnemonic Sounds like "Glue Tonic": Mixing LLMs with Student Feedback to Make Mnemonic Learning Stick
by: Balepur, Nishant, et al.
Published: (2024)
by: Balepur, Nishant, et al.
Published: (2024)
Filling in the Mechanisms: How do LMs Learn Filler-Gap Dependencies under Developmental Constraints?
by: Desai, Atrey, et al.
Published: (2026)
by: Desai, Atrey, et al.
Published: (2026)
Pregnant Questions: The Importance of Pragmatic Awareness in Maternal Health Question Answering
by: Srikanth, Neha, et al.
Published: (2023)
by: Srikanth, Neha, et al.
Published: (2023)
Exploring Design Choices for Building Language-Specific LLMs
by: Tejaswi, Atula, et al.
Published: (2024)
by: Tejaswi, Atula, et al.
Published: (2024)
Can They Dixit? Yes they Can! Dixit as a Playground for Multimodal Language Model Capabilities
by: Balepur, Nishant, et al.
Published: (2025)
by: Balepur, Nishant, et al.
Published: (2025)
Labeled Interactive Topic Models
by: Seelman, Kyle, et al.
Published: (2023)
by: Seelman, Kyle, et al.
Published: (2023)
HistoLens: An Interactive XAI Toolkit for Verifying and Mitigating Flaws in Vision-Language Models for Histopathology
by: Vissapragada, Sandeep, et al.
Published: (2025)
by: Vissapragada, Sandeep, et al.
Published: (2025)
How the Advent of Ubiquitous Large Language Models both Stymie and Turbocharge Dynamic Adversarial Question Generation
by: Sung, Yoo Yeon, et al.
Published: (2024)
by: Sung, Yoo Yeon, et al.
Published: (2024)
At Long Last: The Recognition of Intersectional Discrimination at the ECtHR in FM v Russia
by: Shreya Atrey
Published: (2025)
by: Shreya Atrey
Published: (2025)
Rhapsody: A Dataset for Highlight Detection in Podcasts
by: Park, Younghan, et al.
Published: (2025)
by: Park, Younghan, et al.
Published: (2025)
NAVIG: Natural Language-guided Analysis with Vision Language Models for Image Geo-localization
by: Zhang, Zheyuan, et al.
Published: (2025)
by: Zhang, Zheyuan, et al.
Published: (2025)
NLI under the Microscope: What Atomic Hypothesis Decomposition Reveals
by: Srikanth, Neha, et al.
Published: (2025)
by: Srikanth, Neha, et al.
Published: (2025)
ProxAnn: Use-Oriented Evaluations of Topic Models and Document Clustering
by: Hoyle, Alexander, et al.
Published: (2025)
by: Hoyle, Alexander, et al.
Published: (2025)
Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRA
by: Gor, Maharshi, et al.
Published: (2024)
by: Gor, Maharshi, et al.
Published: (2024)
Evaluating India’s Engagement with the Universal Periodic Review: A Focus on Women’s Rights
by: Gupta, Bhavya
Published: (2024)
by: Gupta, Bhavya
Published: (2024)
Susu Box or Piggy Bank: Assessing Cultural Commonsense Knowledge between Ghana and the U.S
by: Acquaye, Christabel, et al.
Published: (2024)
by: Acquaye, Christabel, et al.
Published: (2024)
On the Influence of Gender and Race in Romantic Relationship Prediction from Large Language Models
by: Sancheti, Abhilasha, et al.
Published: (2024)
by: Sancheti, Abhilasha, et al.
Published: (2024)
CFMatch: Aligning Automated Answer Equivalence Evaluation with Expert Judgments For Open-Domain Question Answering
by: Li, Zongxia, et al.
Published: (2024)
by: Li, Zongxia, et al.
Published: (2024)
How often are errors in natural language reasoning due to paraphrastic variability?
by: Srikanth, Neha, et al.
Published: (2024)
by: Srikanth, Neha, et al.
Published: (2024)
SMART-Editor: A Multi-Agent Framework for Human-Like Design Editing with Structural Integrity
by: Mondal, Ishani, et al.
Published: (2025)
by: Mondal, Ishani, et al.
Published: (2025)
PEDANTS: Cheap but Effective and Interpretable Answer Equivalence
by: Li, Zongxia, et al.
Published: (2024)
by: Li, Zongxia, et al.
Published: (2024)
The Impact of Item-Writing Flaws on Difficulty and Discrimination in Item Response Theory
by: Schmucker, Robin, et al.
Published: (2025)
by: Schmucker, Robin, et al.
Published: (2025)
SciDoc2Diagrammer-MAF: Towards Generation of Scientific Diagrams from Documents guided by Multi-Aspect Feedback Refinement
by: Mondal, Ishani, et al.
Published: (2024)
by: Mondal, Ishani, et al.
Published: (2024)
Is your benchmark truly adversarial? AdvScore: Evaluating Human-Grounded Adversarialness
by: Sung, Yoo Yeon, et al.
Published: (2024)
by: Sung, Yoo Yeon, et al.
Published: (2024)
GRACE: A Granular Benchmark for Evaluating Model Calibration against Human Calibration
by: Sung, Yoo Yeon, et al.
Published: (2025)
by: Sung, Yoo Yeon, et al.
Published: (2025)
On the Mutual Influence of Gender and Occupation in LLM Representations
by: An, Haozhe, et al.
Published: (2025)
by: An, Haozhe, et al.
Published: (2025)
Similar Items
-
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
by: Balepur, Nishant, et al.
Published: (2025) -
Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers
by: Balepur, Nishant, et al.
Published: (2025) -
Is Your Large Language Model Knowledgeable or a Choices-Only Cheater?
by: Balepur, Nishant, et al.
Published: (2024) -
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
by: Balepur, Nishant, et al.
Published: (2024) -
Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can't Answer?
by: Balepur, Nishant, et al.
Published: (2024)