It's Not Easy Being Wrong: Large Language Models Struggle with Process of Elimination Reasoning
Fuente:
arXiv
Guardado en:
| Autores principales: | Balepur, Nishant, Palta, Shramay, Rudinger, Rachel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
por: Palta, Shramay, et al.
Publicado: (2024)
por: Palta, Shramay, et al.
Publicado: (2024)
Is Your Large Language Model Knowledgeable or a Choices-Only Cheater?
por: Balepur, Nishant, et al.
Publicado: (2024)
por: Balepur, Nishant, et al.
Publicado: (2024)
Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers
por: Balepur, Nishant, et al.
Publicado: (2025)
por: Balepur, Nishant, et al.
Publicado: (2025)
Speaking the Right Language: The Impact of Expertise Alignment in User-AI Interactions
por: Palta, Shramay, et al.
Publicado: (2025)
por: Palta, Shramay, et al.
Publicado: (2025)
Everything is Plausible: Investigating the Impact of LLM Rationales on Human Notions of Plausibility
por: Palta, Shramay, et al.
Publicado: (2025)
por: Palta, Shramay, et al.
Publicado: (2025)
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
por: Balepur, Nishant, et al.
Publicado: (2024)
por: Balepur, Nishant, et al.
Publicado: (2024)
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
por: Balepur, Nishant, et al.
Publicado: (2025)
por: Balepur, Nishant, et al.
Publicado: (2025)
Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can't Answer?
por: Balepur, Nishant, et al.
Publicado: (2024)
por: Balepur, Nishant, et al.
Publicado: (2024)
Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User Personas
por: Balepur, Nishant, et al.
Publicado: (2025)
por: Balepur, Nishant, et al.
Publicado: (2025)
Can They Dixit? Yes they Can! Dixit as a Playground for Multimodal Language Model Capabilities
por: Balepur, Nishant, et al.
Publicado: (2025)
por: Balepur, Nishant, et al.
Publicado: (2025)
On the Influence of Gender and Race in Romantic Relationship Prediction from Large Language Models
por: Sancheti, Abhilasha, et al.
Publicado: (2024)
por: Sancheti, Abhilasha, et al.
Publicado: (2024)
A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
por: Balepur, Nishant, et al.
Publicado: (2025)
por: Balepur, Nishant, et al.
Publicado: (2025)
Do Large Language Models Discriminate in Hiring Decisions on the Basis of Race, Ethnicity, and Gender?
por: An, Haozhe, et al.
Publicado: (2024)
por: An, Haozhe, et al.
Publicado: (2024)
BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks
por: Balepur, Nishant, et al.
Publicado: (2026)
por: Balepur, Nishant, et al.
Publicado: (2026)
NLI under the Microscope: What Atomic Hypothesis Decomposition Reveals
por: Srikanth, Neha, et al.
Publicado: (2025)
por: Srikanth, Neha, et al.
Publicado: (2025)
KARL: Knowledge-Aware Retrieval and Representations aid Retention and Learning in Students
por: Shu, Matthew, et al.
Publicado: (2024)
por: Shu, Matthew, et al.
Publicado: (2024)
Language Models Predict Empathy Gaps Between Social In-groups and Out-groups
por: Hou, Yu, et al.
Publicado: (2025)
por: Hou, Yu, et al.
Publicado: (2025)
Easy Problems That LLMs Get Wrong
por: Williams, Sean, et al.
Publicado: (2024)
por: Williams, Sean, et al.
Publicado: (2024)
Susu Box or Piggy Bank: Assessing Cultural Commonsense Knowledge between Ghana and the U.S
por: Acquaye, Christabel, et al.
Publicado: (2024)
por: Acquaye, Christabel, et al.
Publicado: (2024)
EasyInstruct: An Easy-to-use Instruction Processing Framework for Large Language Models
por: Ou, Yixin, et al.
Publicado: (2024)
por: Ou, Yixin, et al.
Publicado: (2024)
DRACULA: Hunting for the Actions Users Want Deep Research Agents to Execute
por: Balepur, Nishant, et al.
Publicado: (2026)
por: Balepur, Nishant, et al.
Publicado: (2026)
'Rich Dad, Poor Lad': How do Large Language Models Contextualize Socioeconomic Factors in College Admission ?
por: Nghiem, Huy, et al.
Publicado: (2025)
por: Nghiem, Huy, et al.
Publicado: (2025)
Constructions Are So Difficult That Even Large Language Models Get Them Right for the Wrong Reasons
por: Zhou, Shijia, et al.
Publicado: (2024)
por: Zhou, Shijia, et al.
Publicado: (2024)
How often are errors in natural language reasoning due to paraphrastic variability?
por: Srikanth, Neha, et al.
Publicado: (2024)
por: Srikanth, Neha, et al.
Publicado: (2024)
Large Language Models Struggle with Unreasonability in Math Problems
por: Ma, Jingyuan, et al.
Publicado: (2024)
por: Ma, Jingyuan, et al.
Publicado: (2024)
DiscoTrace: Representing and Comparing Answering Strategies of Humans and LLMs in Information-Seeking Question Answering
por: Srikanth, Neha, et al.
Publicado: (2026)
por: Srikanth, Neha, et al.
Publicado: (2026)
On the Mutual Influence of Gender and Occupation in LLM Representations
por: An, Haozhe, et al.
Publicado: (2025)
por: An, Haozhe, et al.
Publicado: (2025)
Natural Language Inference Improves Compositionality in Vision-Language Models
por: Cascante-Bonilla, Paola, et al.
Publicado: (2024)
por: Cascante-Bonilla, Paola, et al.
Publicado: (2024)
Language Models Still Struggle to Zero-shot Reason about Time Series
por: Merrill, Mike A., et al.
Publicado: (2024)
por: Merrill, Mike A., et al.
Publicado: (2024)
Large Reasoning Models Struggle to Transfer Parametric Knowledge Across Scripts
por: Bandarkar, Lucas, et al.
Publicado: (2026)
por: Bandarkar, Lucas, et al.
Publicado: (2026)
Wrong-of-Thought: An Integrated Reasoning Framework with Multi-Perspective Verification and Wrong Information
por: Zhang, Yongheng, et al.
Publicado: (2024)
por: Zhang, Yongheng, et al.
Publicado: (2024)
Why Do Large Language Models (LLMs) Struggle to Count Letters?
por: Fu, Tairan, et al.
Publicado: (2024)
por: Fu, Tairan, et al.
Publicado: (2024)
Exploring Large Language Models to generate Easy to Read content
por: Martínez, Paloma, et al.
Publicado: (2024)
por: Martínez, Paloma, et al.
Publicado: (2024)
Generics and Default Reasoning in Large Language Models
por: Kirkpatrick, James Ravi, et al.
Publicado: (2025)
por: Kirkpatrick, James Ravi, et al.
Publicado: (2025)
EEVEE: An Easy Annotation Tool for Natural Language Processing
por: Sorensen, Axel, et al.
Publicado: (2024)
por: Sorensen, Axel, et al.
Publicado: (2024)
Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong
por: Fu, Tairan, et al.
Publicado: (2025)
por: Fu, Tairan, et al.
Publicado: (2025)
FRIDA to the Rescue! Analyzing Synthetic Data Effectiveness in Object-Based Common Sense Reasoning for Disaster Response
por: Shichman, Mollie, et al.
Publicado: (2025)
por: Shichman, Mollie, et al.
Publicado: (2025)
Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users
por: Balepur, Nishant, et al.
Publicado: (2026)
por: Balepur, Nishant, et al.
Publicado: (2026)
Take Out Your Calculators: Estimating the Real Difficulty of Question Items with LLM Student Simulations
por: Acquaye, Christabel, et al.
Publicado: (2026)
por: Acquaye, Christabel, et al.
Publicado: (2026)
All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Language
por: Guo, Shiyuan, et al.
Publicado: (2025)
por: Guo, Shiyuan, et al.
Publicado: (2025)
Ejemplares similares
-
Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
por: Palta, Shramay, et al.
Publicado: (2024) -
Is Your Large Language Model Knowledgeable or a Choices-Only Cheater?
por: Balepur, Nishant, et al.
Publicado: (2024) -
Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers
por: Balepur, Nishant, et al.
Publicado: (2025) -
Speaking the Right Language: The Impact of Expertise Alignment in User-AI Interactions
por: Palta, Shramay, et al.
Publicado: (2025) -
Everything is Plausible: Investigating the Impact of LLM Rationales on Human Notions of Plausibility
por: Palta, Shramay, et al.
Publicado: (2025)