Wait, that's not an option: LLMs Robustness with Incorrect Multiple-Choice Options
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Góral, Gracjan, Wiśnios, Emilia, Sankowski, Piotr, Budzianowski, Paweł |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
OpenGVL -- Benchmarking Visual Temporal Progress for Data Curation
par: Budzianowski, Paweł, et autres
Publié: (2025)
par: Budzianowski, Paweł, et autres
Publié: (2025)
LLMs May Perform MCQA by Selecting the Least Incorrect Option
par: Wang, Haochun, et autres
Publié: (2024)
par: Wang, Haochun, et autres
Publié: (2024)
Pheme: Efficient and Conversational Speech Generation
par: Budzianowski, Paweł, et autres
Publié: (2024)
par: Budzianowski, Paweł, et autres
Publié: (2024)
Option-ID Based Elimination For Multiple Choice Questions
par: Zhu, Zhenhao, et autres
Publié: (2025)
par: Zhu, Zhenhao, et autres
Publié: (2025)
It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education
par: Singh, Shrutika, et autres
Publié: (2025)
par: Singh, Shrutika, et autres
Publié: (2025)
Beyond Recognition: Evaluating Visual Perspective Taking in Vision Language Models
par: Góral, Gracjan, et autres
Publié: (2025)
par: Góral, Gracjan, et autres
Publié: (2025)
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
par: Chanin, David, et autres
Publié: (2025)
par: Chanin, David, et autres
Publié: (2025)
Beyond Multiple Choice: Verifiable OpenQA for Robust Vision-Language RFT
par: Liu, Yesheng, et autres
Publié: (2025)
par: Liu, Yesheng, et autres
Publié: (2025)
Can LLMs Model Incorrect Student Reasoning? A Case Study on Distractor Generation
par: Zengaffinen, Yanick, et autres
Publié: (2026)
par: Zengaffinen, Yanick, et autres
Publié: (2026)
Differentiating Choices via Commonality for Multiple-Choice Question Answering
par: Deng, Wenqing, et autres
Publié: (2024)
par: Deng, Wenqing, et autres
Publié: (2024)
Can We Verify Step by Step for Incorrect Answer Detection?
par: Xu, Xin, et autres
Publié: (2024)
par: Xu, Xin, et autres
Publié: (2024)
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects
par: Mustapha, Ahmad, et autres
Publié: (2024)
par: Mustapha, Ahmad, et autres
Publié: (2024)
The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks
par: Pinhanez, Claudio, et autres
Publié: (2025)
par: Pinhanez, Claudio, et autres
Publié: (2025)
Improving Score Reliability of Multiple Choice Benchmarks with Consistency Evaluation and Altered Answer Choices
par: Cavalin, Paulo, et autres
Publié: (2025)
par: Cavalin, Paulo, et autres
Publié: (2025)
Look at the Text: Instruction-Tuned Language Models are More Robust Multiple Choice Selectors than You Think
par: Wang, Xinpeng, et autres
Publié: (2024)
par: Wang, Xinpeng, et autres
Publié: (2024)
Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong
par: Fu, Tairan, et autres
Publié: (2025)
par: Fu, Tairan, et autres
Publié: (2025)
Seeing Through Their Eyes: Evaluating Visual Perspective Taking in Vision Language Models
par: Góral, Gracjan, et autres
Publié: (2024)
par: Góral, Gracjan, et autres
Publié: (2024)
Question Difficulty Ranking for Multiple-Choice Reading Comprehension
par: Raina, Vatsal, et autres
Publié: (2024)
par: Raina, Vatsal, et autres
Publié: (2024)
Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A
par: Plaut, Benjamin, et autres
Publié: (2024)
par: Plaut, Benjamin, et autres
Publié: (2024)
Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
par: Palta, Shramay, et autres
Publié: (2024)
par: Palta, Shramay, et autres
Publié: (2024)
Large-Scale Aspect-Based Sentiment Analysis with Reasoning-Infused LLMs
par: Liskowski, Paweł, et autres
Publié: (2026)
par: Liskowski, Paweł, et autres
Publié: (2026)
Automated Generation and Tagging of Knowledge Components from Multiple-Choice Questions
par: Moore, Steven, et autres
Publié: (2024)
par: Moore, Steven, et autres
Publié: (2024)
A Vietnamese Dataset for Text Segmentation and Multiple Choices Reading Comprehension
par: Hai, Toan Nguyen, et autres
Publié: (2025)
par: Hai, Toan Nguyen, et autres
Publié: (2025)
Self-Correcting Large Language Models: Generation vs. Multiple Choice
par: Rahmani, Hossein A., et autres
Publié: (2025)
par: Rahmani, Hossein A., et autres
Publié: (2025)
UnibucLLM: Harnessing LLMs for Automated Prediction of Item Difficulty and Response Time for Multiple-Choice Questions
par: Rogoz, Ana-Cristina, et autres
Publié: (2024)
par: Rogoz, Ana-Cristina, et autres
Publié: (2024)
Generating Plausible Distractors for Multiple-Choice Questions via Student Choice Prediction
par: Lee, Yooseop, et autres
Publié: (2025)
par: Lee, Yooseop, et autres
Publié: (2025)
Stepwise Guided Policy Optimization: Coloring your Incorrect Reasoning in GRPO
par: Chen, Peter, et autres
Publié: (2025)
par: Chen, Peter, et autres
Publié: (2025)
From Multiple-Choice to Extractive QA: A Case Study for English and Arabic
par: Lynn, Teresa, et autres
Publié: (2024)
par: Lynn, Teresa, et autres
Publié: (2024)
Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
par: Wiegreffe, Sarah, et autres
Publié: (2024)
par: Wiegreffe, Sarah, et autres
Publié: (2024)
Boosting Process-Correct CoT Reasoning by Modeling Solvability of Multiple-Choice QA
par: Schumann, Raphael, et autres
Publié: (2025)
par: Schumann, Raphael, et autres
Publié: (2025)
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
par: Ludziejewski, Jan, et autres
Publié: (2025)
par: Ludziejewski, Jan, et autres
Publié: (2025)
Multiple-Choice Question Generation Using Large Language Models: Methodology and Educator Insights
par: Biancini, Giorgio, et autres
Publié: (2025)
par: Biancini, Giorgio, et autres
Publié: (2025)
Conformal P-Value in Multiple-Choice Question Answering Tasks with Provable Risk Control
par: Ye, Yuanchang
Publié: (2025)
par: Ye, Yuanchang
Publié: (2025)
GeoChallenge: A Multi-Answer Multiple-Choice Benchmark for Geometric Reasoning with Diagrams
par: Zhang, Yushun, et autres
Publié: (2026)
par: Zhang, Yushun, et autres
Publié: (2026)
How Can I Get It Right? Using GPT to Rephrase Incorrect Trainee Responses
par: Lin, Jionghao, et autres
Publié: (2024)
par: Lin, Jionghao, et autres
Publié: (2024)
Evaluating LLMs with Multiple Problems at once
par: Wang, Zhengxiang, et autres
Publié: (2024)
par: Wang, Zhengxiang, et autres
Publié: (2024)
Cannot See the Forest for the Trees: Invoking Heuristics and Biases to Elicit Irrational Choices of LLMs
par: Yang, Haoming, et autres
Publié: (2025)
par: Yang, Haoming, et autres
Publié: (2025)
Biomedical Entity Linking as Multiple Choice Question Answering
par: Lin, Zhenxi, et autres
Publié: (2024)
par: Lin, Zhenxi, et autres
Publié: (2024)
Balancing Rigor and Utility: Mitigating Cognitive Biases in Large Language Models for Multiple-Choice Questions
par: Zhong, Hanyang, et autres
Publié: (2024)
par: Zhong, Hanyang, et autres
Publié: (2024)
Exploring Iterative Enhancement for Improving Learnersourced Multiple-Choice Question Explanations with Large Language Models
par: Bao, Qiming, et autres
Publié: (2023)
par: Bao, Qiming, et autres
Publié: (2023)
Documents similaires
-
OpenGVL -- Benchmarking Visual Temporal Progress for Data Curation
par: Budzianowski, Paweł, et autres
Publié: (2025) -
LLMs May Perform MCQA by Selecting the Least Incorrect Option
par: Wang, Haochun, et autres
Publié: (2024) -
Pheme: Efficient and Conversational Speech Generation
par: Budzianowski, Paweł, et autres
Publié: (2024) -
Option-ID Based Elimination For Multiple Choice Questions
par: Zhu, Zhenhao, et autres
Publié: (2025) -
It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education
par: Singh, Shrutika, et autres
Publié: (2025)