Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | Palta, Shramay, Balepur, Nishant, Rankel, Peter, Wiegreffe, Sarah, Carpuat, Marine, Rudinger, Rachel |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Everything is Plausible: Investigating the Impact of LLM Rationales on Human Notions of Plausibility
di: Palta, Shramay, et al.
Pubblicazione: (2025)
di: Palta, Shramay, et al.
Pubblicazione: (2025)
It's Not Easy Being Wrong: Large Language Models Struggle with Process of Elimination Reasoning
di: Balepur, Nishant, et al.
Pubblicazione: (2023)
di: Balepur, Nishant, et al.
Pubblicazione: (2023)
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
di: Balepur, Nishant, et al.
Pubblicazione: (2024)
di: Balepur, Nishant, et al.
Pubblicazione: (2024)
Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
Speaking the Right Language: The Impact of Expertise Alignment in User-AI Interactions
di: Palta, Shramay, et al.
Pubblicazione: (2025)
di: Palta, Shramay, et al.
Pubblicazione: (2025)
Multiple LLM Agents Debate for Equitable Cultural Alignment
di: Ki, Dayeon, et al.
Pubblicazione: (2025)
di: Ki, Dayeon, et al.
Pubblicazione: (2025)
Is Your Large Language Model Knowledgeable or a Choices-Only Cheater?
di: Balepur, Nishant, et al.
Pubblicazione: (2024)
di: Balepur, Nishant, et al.
Pubblicazione: (2024)
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
di: Wiegreffe, Sarah, et al.
Pubblicazione: (2024)
di: Wiegreffe, Sarah, et al.
Pubblicazione: (2024)
Should We be Pedantic About Reasoning Errors in Machine Translation?
di: Bao, Calvin, et al.
Pubblicazione: (2026)
di: Bao, Calvin, et al.
Pubblicazione: (2026)
Reheat Nachos for Dinner? Evaluating AI Support for Cross-Cultural Communication of Neologisms
di: Ki, Dayeon, et al.
Pubblicazione: (2026)
di: Ki, Dayeon, et al.
Pubblicazione: (2026)
What Makes Good Multilingual Reasoning? Disentangling Reasoning Traces with Measurable Features
di: Ki, Dayeon, et al.
Pubblicazione: (2026)
di: Ki, Dayeon, et al.
Pubblicazione: (2026)
Generating Plausible Distractors for Multiple-Choice Questions via Student Choice Prediction
di: Lee, Yooseop, et al.
Pubblicazione: (2025)
di: Lee, Yooseop, et al.
Pubblicazione: (2025)
How often are errors in natural language reasoning due to paraphrastic variability?
di: Srikanth, Neha, et al.
Pubblicazione: (2024)
di: Srikanth, Neha, et al.
Pubblicazione: (2024)
Guiding Large Language Models to Post-Edit Machine Translation with Error Annotations
di: Ki, Dayeon, et al.
Pubblicazione: (2024)
di: Ki, Dayeon, et al.
Pubblicazione: (2024)
Keep It Private: Unsupervised Privatization of Online Text
di: Bao, Calvin, et al.
Pubblicazione: (2024)
di: Bao, Calvin, et al.
Pubblicazione: (2024)
Automatic Input Rewriting Improves Translation with Large Language Models
di: Ki, Dayeon, et al.
Pubblicazione: (2025)
di: Ki, Dayeon, et al.
Pubblicazione: (2025)
Plausibility as Commonsense Reasoning: Humans Succeed, Large Language Models Do not
di: Karakaş, Sercan
Pubblicazione: (2026)
di: Karakaş, Sercan
Pubblicazione: (2026)
Take Out Your Calculators: Estimating the Real Difficulty of Question Items with LLM Student Simulations
di: Acquaye, Christabel, et al.
Pubblicazione: (2026)
di: Acquaye, Christabel, et al.
Pubblicazione: (2026)
Can They Dixit? Yes they Can! Dixit as a Playground for Multimodal Language Model Capabilities
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
Estimating Commonsense Plausibility through Semantic Shifts
di: Cui, Wanqing, et al.
Pubblicazione: (2025)
di: Cui, Wanqing, et al.
Pubblicazione: (2025)
Can you map it to English? The Role of Cross-Lingual Alignment in Multilingual Performance of LLMs
di: Ravisankar, Kartik, et al.
Pubblicazione: (2025)
di: Ravisankar, Kartik, et al.
Pubblicazione: (2025)
Should I Share this Translation? Evaluating Quality Feedback for User Reliance on Machine Translation
di: Ki, Dayeon, et al.
Pubblicazione: (2025)
di: Ki, Dayeon, et al.
Pubblicazione: (2025)
Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can't Answer?
di: Balepur, Nishant, et al.
Pubblicazione: (2024)
di: Balepur, Nishant, et al.
Pubblicazione: (2024)
BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks
di: Balepur, Nishant, et al.
Pubblicazione: (2026)
di: Balepur, Nishant, et al.
Pubblicazione: (2026)
GraphicBench: A Planning Benchmark for Graphic Design with Language Agents
di: Ki, Dayeon, et al.
Pubblicazione: (2025)
di: Ki, Dayeon, et al.
Pubblicazione: (2025)
LOGICAL-COMMONSENSEQA: A Benchmark for Logical Commonsense Reasoning
di: Junias, Obed, et al.
Pubblicazione: (2026)
di: Junias, Obed, et al.
Pubblicazione: (2026)
PerCoR: Evaluating Commonsense Reasoning in Persian via Multiple-Choice Sentence Completion
di: Alikhani, Morteza, et al.
Pubblicazione: (2025)
di: Alikhani, Morteza, et al.
Pubblicazione: (2025)
Mechanistic?
di: Saphra, Naomi, et al.
Pubblicazione: (2024)
di: Saphra, Naomi, et al.
Pubblicazione: (2024)
Benchmarking Chinese Commonsense Reasoning with a Multi-hop Reasoning Perspective
di: You, Wangjie, et al.
Pubblicazione: (2025)
di: You, Wangjie, et al.
Pubblicazione: (2025)
Differentiating Choices via Commonality for Multiple-Choice Question Answering
di: Deng, Wenqing, et al.
Pubblicazione: (2024)
di: Deng, Wenqing, et al.
Pubblicazione: (2024)
Focus on Your Question! Interpreting and Mitigating Toxic CoT Problems in Commonsense Reasoning
di: Li, Jiachun, et al.
Pubblicazione: (2024)
di: Li, Jiachun, et al.
Pubblicazione: (2024)
GRASP: A Grid-Based Benchmark for Evaluating Commonsense Spatial Reasoning
di: Tang, Zhisheng, et al.
Pubblicazione: (2024)
di: Tang, Zhisheng, et al.
Pubblicazione: (2024)
The Unreasonable Effectiveness of Easy Training Data for Hard Tasks
di: Hase, Peter, et al.
Pubblicazione: (2024)
di: Hase, Peter, et al.
Pubblicazione: (2024)
Question Difficulty Ranking for Multiple-Choice Reading Comprehension
di: Raina, Vatsal, et al.
Pubblicazione: (2024)
di: Raina, Vatsal, et al.
Pubblicazione: (2024)
GeoChallenge: A Multi-Answer Multiple-Choice Benchmark for Geometric Reasoning with Diagrams
di: Zhang, Yushun, et al.
Pubblicazione: (2026)
di: Zhang, Yushun, et al.
Pubblicazione: (2026)
Multimodal Commonsense Knowledge Distillation for Visual Question Answering
di: Yang, Shuo, et al.
Pubblicazione: (2024)
di: Yang, Shuo, et al.
Pubblicazione: (2024)
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects
di: Mustapha, Ahmad, et al.
Pubblicazione: (2024)
di: Mustapha, Ahmad, et al.
Pubblicazione: (2024)
Language Models Predict Empathy Gaps Between Social In-groups and Out-groups
di: Hou, Yu, et al.
Pubblicazione: (2025)
di: Hou, Yu, et al.
Pubblicazione: (2025)
What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
di: Cheng, Stephen, et al.
Pubblicazione: (2026)
di: Cheng, Stephen, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Everything is Plausible: Investigating the Impact of LLM Rationales on Human Notions of Plausibility
di: Palta, Shramay, et al.
Pubblicazione: (2025) -
It's Not Easy Being Wrong: Large Language Models Struggle with Process of Elimination Reasoning
di: Balepur, Nishant, et al.
Pubblicazione: (2023) -
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
di: Balepur, Nishant, et al.
Pubblicazione: (2024) -
Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers
di: Balepur, Nishant, et al.
Pubblicazione: (2025) -
Speaking the Right Language: The Impact of Expertise Alignment in User-AI Interactions
di: Palta, Shramay, et al.
Pubblicazione: (2025)