Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
Fuente:
arXiv
Saved in:
| Main Authors: | Balepur, Nishant, Ravichander, Abhilasha, Rudinger, Rachel |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can't Answer?
by: Balepur, Nishant, et al.
Published: (2024)
by: Balepur, Nishant, et al.
Published: (2024)
Is Your Large Language Model Knowledgeable or a Choices-Only Cheater?
by: Balepur, Nishant, et al.
Published: (2024)
by: Balepur, Nishant, et al.
Published: (2024)
Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers
by: Balepur, Nishant, et al.
Published: (2025)
by: Balepur, Nishant, et al.
Published: (2025)
Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
by: Palta, Shramay, et al.
Published: (2024)
by: Palta, Shramay, et al.
Published: (2024)
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
by: Balepur, Nishant, et al.
Published: (2025)
by: Balepur, Nishant, et al.
Published: (2025)
It's Not Easy Being Wrong: Large Language Models Struggle with Process of Elimination Reasoning
by: Balepur, Nishant, et al.
Published: (2023)
by: Balepur, Nishant, et al.
Published: (2023)
Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
by: Wiegreffe, Sarah, et al.
Published: (2024)
by: Wiegreffe, Sarah, et al.
Published: (2024)
On the Influence of Gender and Race in Romantic Relationship Prediction from Large Language Models
by: Sancheti, Abhilasha, et al.
Published: (2024)
by: Sancheti, Abhilasha, et al.
Published: (2024)
BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks
by: Balepur, Nishant, et al.
Published: (2026)
by: Balepur, Nishant, et al.
Published: (2026)
DiscoTrace: Representing and Comparing Answering Strategies of Humans and LLMs in Information-Seeking Question Answering
by: Srikanth, Neha, et al.
Published: (2026)
by: Srikanth, Neha, et al.
Published: (2026)
Do Large Language Models Plan Answer Positions? Position Bias in Multiple-Choice Question Generation
by: Tang, Xuemei, et al.
Published: (2026)
by: Tang, Xuemei, et al.
Published: (2026)
Polyglots or Multitudes? Multilingual LLM Answers to Value-laden Multiple-Choice Questions
by: Labat, Léo, et al.
Published: (2026)
by: Labat, Léo, et al.
Published: (2026)
On the Mutual Influence of Gender and Occupation in LLM Representations
by: An, Haozhe, et al.
Published: (2025)
by: An, Haozhe, et al.
Published: (2025)
Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User Personas
by: Balepur, Nishant, et al.
Published: (2025)
by: Balepur, Nishant, et al.
Published: (2025)
Why and How LLMs Hallucinate: Connecting the Dots with Subsequence Associations
by: Sun, Yiyou, et al.
Published: (2025)
by: Sun, Yiyou, et al.
Published: (2025)
Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions
by: Park, Yoonah, et al.
Published: (2025)
by: Park, Yoonah, et al.
Published: (2025)
Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering
by: Molfese, Francesco Maria, et al.
Published: (2025)
by: Molfese, Francesco Maria, et al.
Published: (2025)
Anchored Answers: Unravelling Positional Bias in GPT-2's Multiple-Choice Questions
by: Li, Ruizhe, et al.
Published: (2024)
by: Li, Ruizhe, et al.
Published: (2024)
What Has Been Lost with Synthetic Evaluation?
by: Gill, Alexander, et al.
Published: (2025)
by: Gill, Alexander, et al.
Published: (2025)
LLMs Provide Unstable Answers to Legal Questions
by: Blair-Stanek, Andrew, et al.
Published: (2025)
by: Blair-Stanek, Andrew, et al.
Published: (2025)
Pregnant Questions: The Importance of Pragmatic Awareness in Maternal Health Question Answering
by: Srikanth, Neha, et al.
Published: (2023)
by: Srikanth, Neha, et al.
Published: (2023)
From Answers to Questions: EQGBench for Evaluating LLMs' Educational Question Generation
by: Zhou, Chengliang, et al.
Published: (2025)
by: Zhou, Chengliang, et al.
Published: (2025)
Take Out Your Calculators: Estimating the Real Difficulty of Question Items with LLM Student Simulations
by: Acquaye, Christabel, et al.
Published: (2026)
by: Acquaye, Christabel, et al.
Published: (2026)
Question: How do Large Language Models perform on the Question Answering tasks? Answer:
by: Fischer, Kevin, et al.
Published: (2024)
by: Fischer, Kevin, et al.
Published: (2024)
Mind the Gap: A Closer Look at Tokenization for Multiple-Choice Question Answering with LLMs
by: Sanz-Guerrero, Mario, et al.
Published: (2025)
by: Sanz-Guerrero, Mario, et al.
Published: (2025)
Do LLMs Provide Consistent Answers to Health-Related Questions across Languages?
by: Schlicht, Ipek Baris, et al.
Published: (2025)
by: Schlicht, Ipek Baris, et al.
Published: (2025)
Interpreting Answers to Yes-No Questions in Dialogues from Multiple Domains
by: Wang, Zijie, et al.
Published: (2024)
by: Wang, Zijie, et al.
Published: (2024)
Evaluating and Calibrating LLM Confidence on Questions with Multiple Correct Answers
by: Wang, Yuhan, et al.
Published: (2026)
by: Wang, Yuhan, et al.
Published: (2026)
Differentiating Choices via Commonality for Multiple-Choice Question Answering
by: Deng, Wenqing, et al.
Published: (2024)
by: Deng, Wenqing, et al.
Published: (2024)
Multiple-Choice Questions are Efficient and Robust LLM Evaluators
by: Zhang, Ziyin, et al.
Published: (2024)
by: Zhang, Ziyin, et al.
Published: (2024)
HALoGEN: Fantastic LLM Hallucinations and Where to Find Them
by: Ravichander, Abhilasha, et al.
Published: (2025)
by: Ravichander, Abhilasha, et al.
Published: (2025)
Question Answering with LLMs and Learning from Answer Sets
by: Borroto, Manuel, et al.
Published: (2025)
by: Borroto, Manuel, et al.
Published: (2025)
Knowledge-Augmented Question Error Correction for Chinese Question Answer System with QuestionRAG
by: Qiu, Longpeng, et al.
Published: (2025)
by: Qiu, Longpeng, et al.
Published: (2025)
A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
by: Balepur, Nishant, et al.
Published: (2025)
by: Balepur, Nishant, et al.
Published: (2025)
Can They Dixit? Yes they Can! Dixit as a Playground for Multimodal Language Model Capabilities
by: Balepur, Nishant, et al.
Published: (2025)
by: Balepur, Nishant, et al.
Published: (2025)
Question Difficulty Ranking for Multiple-Choice Reading Comprehension
by: Raina, Vatsal, et al.
Published: (2024)
by: Raina, Vatsal, et al.
Published: (2024)
Do LLMs Find Human Answers To Fact-Driven Questions Perplexing? A Case Study on Reddit
by: Seegmiller, Parker, et al.
Published: (2024)
by: Seegmiller, Parker, et al.
Published: (2024)
UBench: Benchmarking Uncertainty in Large Language Models with Multiple Choice Questions
by: Wang, Xunzhi, et al.
Published: (2024)
by: Wang, Xunzhi, et al.
Published: (2024)
LLM Distillation for Efficient Few-Shot Multiple Choice Question Answering
by: Sutanto, Patrick, et al.
Published: (2024)
by: Sutanto, Patrick, et al.
Published: (2024)
How often are errors in natural language reasoning due to paraphrastic variability?
by: Srikanth, Neha, et al.
Published: (2024)
by: Srikanth, Neha, et al.
Published: (2024)
Similar Items
-
Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can't Answer?
by: Balepur, Nishant, et al.
Published: (2024) -
Is Your Large Language Model Knowledgeable or a Choices-Only Cheater?
by: Balepur, Nishant, et al.
Published: (2024) -
Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers
by: Balepur, Nishant, et al.
Published: (2025) -
Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
by: Palta, Shramay, et al.
Published: (2024) -
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
by: Balepur, Nishant, et al.
Published: (2025)