Position: On the Methodological Pitfalls of Evaluating Base LLMs for Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Chan, Jason, Zhao, Zhixue, Gaizauskas, Robert |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Position: Logical Soundness is not a Reliable Criterion for Neurosymbolic Fact-Checking with LLMs
by: Chan, Jason, et al.
Published: (2026)
by: Chan, Jason, et al.
Published: (2026)
RULEBREAKERS: Challenging LLMs at the Crossroads between Formal Logic and Human-like Reasoning
by: Chan, Jason, et al.
Published: (2024)
by: Chan, Jason, et al.
Published: (2024)
Explanation Generation for Contradiction Reconciliation with LLMs
by: Chan, Jason, et al.
Published: (2026)
by: Chan, Jason, et al.
Published: (2026)
When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs
by: Li, Xiaomin, et al.
Published: (2025)
by: Li, Xiaomin, et al.
Published: (2025)
It's All About In-Context Learning! Teaching Extremely Low-Resource Languages to LLMs
by: Li, Yue, et al.
Published: (2025)
by: Li, Yue, et al.
Published: (2025)
Tracing and Reversing Edits in LLMs
by: Youssef, Paul, et al.
Published: (2025)
by: Youssef, Paul, et al.
Published: (2025)
Evaluating the Promise and Pitfalls of LLMs in Hiring Decisions
by: Anzenberg, Eitan, et al.
Published: (2025)
by: Anzenberg, Eitan, et al.
Published: (2025)
Right Is Not Enough: The Pitfalls of Outcome Supervision in Training LLMs for Math Reasoning
by: Guo, Jiaxing, et al.
Published: (2025)
by: Guo, Jiaxing, et al.
Published: (2025)
How to Make LLMs Forget: On Reversing In-Context Knowledge Edits
by: Youssef, Paul, et al.
Published: (2024)
by: Youssef, Paul, et al.
Published: (2024)
Can LLMs $\textit{understand}$ Math? -- Exploring the Pitfalls in Mathematical Reasoning
by: Roy, Tiasa Singha, et al.
Published: (2025)
by: Roy, Tiasa Singha, et al.
Published: (2025)
Position: Editing Large Language Models Poses Serious Safety Risks
by: Youssef, Paul, et al.
Published: (2025)
by: Youssef, Paul, et al.
Published: (2025)
Can Confidence Estimates Decide When Chain-of-Thought Is Necessary for LLMs?
by: Lewis-Lim, Samuel, et al.
Published: (2025)
by: Lewis-Lim, Samuel, et al.
Published: (2025)
From Early Encoding to Late Suppression: Interpreting LLMs on Character Counting Tasks
by: Datta, Ayan, et al.
Published: (2026)
by: Datta, Ayan, et al.
Published: (2026)
AraReasoner: Evaluating Reasoning-Based LLMs for Arabic NLP
by: Hasanaath, Ahmed, et al.
Published: (2025)
by: Hasanaath, Ahmed, et al.
Published: (2025)
Pitfalls of Conversational LLMs on News Debiasing
by: Schlicht, Ipek Baris, et al.
Published: (2024)
by: Schlicht, Ipek Baris, et al.
Published: (2024)
Comparing Explanation Faithfulness between Multilingual and Monolingual Fine-tuned Language Models
by: Zhao, Zhixue, et al.
Published: (2024)
by: Zhao, Zhixue, et al.
Published: (2024)
Beware of Reasoning Overconfidence: Pitfalls in the Reasoning Process for Multi-solution Tasks
by: Guan, Jiannan, et al.
Published: (2025)
by: Guan, Jiannan, et al.
Published: (2025)
Survey-to-Behavior: Downstream Alignment of Human Values in LLMs via Survey Questions
by: Nie, Shangrui, et al.
Published: (2025)
by: Nie, Shangrui, et al.
Published: (2025)
FINEREASON: Evaluating and Improving LLMs' Deliberate Reasoning through Reflective Puzzle Solving
by: Chen, Guizhen, et al.
Published: (2025)
by: Chen, Guizhen, et al.
Published: (2025)
Training and Evaluation of Guideline-Based Medical Reasoning in LLMs
by: Staniek, Michael, et al.
Published: (2025)
by: Staniek, Michael, et al.
Published: (2025)
Label Set Optimization via Activation Distribution Kurtosis for Zero-shot Classification with Generative Models
by: Li, Yue, et al.
Published: (2024)
by: Li, Yue, et al.
Published: (2024)
Pitfalls of Evaluating Language Models with Open Benchmarks
by: Hasan, Md. Najib, et al.
Published: (2025)
by: Hasan, Md. Najib, et al.
Published: (2025)
Incorporating Attribution Importance for Improving Faithfulness Metrics
by: Zhao, Zhixue, et al.
Published: (2023)
by: Zhao, Zhixue, et al.
Published: (2023)
ReAGent: A Model-agnostic Feature Attribution Method for Generative Language Models
by: Zhao, Zhixue, et al.
Published: (2024)
by: Zhao, Zhixue, et al.
Published: (2024)
Mitigating Content Effects on Reasoning in Language Models through Fine-Grained Activation Steering
by: Valentino, Marco, et al.
Published: (2025)
by: Valentino, Marco, et al.
Published: (2025)
Disentangling Mathematical Reasoning in LLMs: A Methodological Investigation of Internal Mechanisms
by: Baeumel, Tanja, et al.
Published: (2026)
by: Baeumel, Tanja, et al.
Published: (2026)
The Pitfalls of Growing Group Complexity: LLMs and Social Choice-Based Aggregation for Group Recommendations
by: Waterschoot, Cedric, et al.
Published: (2025)
by: Waterschoot, Cedric, et al.
Published: (2025)
Exploring Vision Language Models for Multimodal and Multilingual Stance Detection
by: Vasilakes, Jake, et al.
Published: (2025)
by: Vasilakes, Jake, et al.
Published: (2025)
Two Failures of Self-Consistency in the Multi-Step Reasoning of LLMs
by: Chen, Angelica, et al.
Published: (2023)
by: Chen, Angelica, et al.
Published: (2023)
SCRum-9: Multilingual Stance Classification over Rumours on Social Media
by: Li, Yue, et al.
Published: (2025)
by: Li, Yue, et al.
Published: (2025)
Do LLMs Provide Consistent Answers to Health-Related Questions across Languages?
by: Schlicht, Ipek Baris, et al.
Published: (2025)
by: Schlicht, Ipek Baris, et al.
Published: (2025)
Efficient Pruning of Text-to-Image Models: Insights from Pruning Stable Diffusion
by: Ramesh, Samarth N, et al.
Published: (2024)
by: Ramesh, Samarth N, et al.
Published: (2024)
The Pitfalls of Defining Hallucination
by: van Deemter, Kees
Published: (2024)
by: van Deemter, Kees
Published: (2024)
PPA-Plan: Proactive Pitfall Avoidance for Reliable Planning in Long-Context LLM Reasoning
by: Kim, Byeongjin, et al.
Published: (2026)
by: Kim, Byeongjin, et al.
Published: (2026)
Don't Get Lost in the Trees: Streamlining LLM Reasoning by Overcoming Tree Search Exploration Pitfalls
by: Wang, Ante, et al.
Published: (2025)
by: Wang, Ante, et al.
Published: (2025)
Evaluating Hierarchical Clinical Document Classification Using Reasoning-Based LLMs
by: Mustafa, Akram, et al.
Published: (2025)
by: Mustafa, Akram, et al.
Published: (2025)
Pitfalls and Outlooks in Using COMET
by: Zouhar, Vilém, et al.
Published: (2024)
by: Zouhar, Vilém, et al.
Published: (2024)
A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility
by: Hochlehnert, Andreas, et al.
Published: (2025)
by: Hochlehnert, Andreas, et al.
Published: (2025)
Evaluating o1-Like LLMs: Unlocking Reasoning for Translation through Comprehensive Analysis
by: Chen, Andong, et al.
Published: (2025)
by: Chen, Andong, et al.
Published: (2025)
A 2-step Framework for Automated Literary Translation Evaluation: Its Promises and Pitfalls
by: Shafayat, Sheikh, et al.
Published: (2024)
by: Shafayat, Sheikh, et al.
Published: (2024)
Similar Items
-
Position: Logical Soundness is not a Reliable Criterion for Neurosymbolic Fact-Checking with LLMs
by: Chan, Jason, et al.
Published: (2026) -
RULEBREAKERS: Challenging LLMs at the Crossroads between Formal Logic and Human-like Reasoning
by: Chan, Jason, et al.
Published: (2024) -
Explanation Generation for Contradiction Reconciliation with LLMs
by: Chan, Jason, et al.
Published: (2026) -
When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs
by: Li, Xiaomin, et al.
Published: (2025) -
It's All About In-Context Learning! Teaching Extremely Low-Resource Languages to LLMs
by: Li, Yue, et al.
Published: (2025)