What Defines Good Reasoning in LLMs? Dissecting Reasoning Steps with Multi-Aspect Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Do, Heejin, Hwang, Jaehui, Han, Dongyoon, Oh, Seong Joon, Yun, Sangdoo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Oops, Wait: Token-Level Signals as a Lens into LLM Reasoning
by: Hwang, Jaehui, et al.
Published: (2026)
by: Hwang, Jaehui, et al.
Published: (2026)
Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers
by: Green, Tommaso, et al.
Published: (2025)
by: Green, Tommaso, et al.
Published: (2025)
Dr.LLM: Dynamic Layer Routing in LLMs
by: Heakl, Ahmed, et al.
Published: (2025)
by: Heakl, Ahmed, et al.
Published: (2025)
Scaling Up Membership Inference: When and How Attacks Succeed on Large Language Models
by: Puerto, Haritz, et al.
Published: (2024)
by: Puerto, Haritz, et al.
Published: (2024)
StoryCoder: Narrative Reformulation for Structured Reasoning in LLM Code Generation
by: Jang, Geonhui, et al.
Published: (2026)
by: Jang, Geonhui, et al.
Published: (2026)
Teach-to-Reason with Scoring: Self-Explainable Rationale-Driven Multi-Trait Essay Scoring
by: Do, Heejin, et al.
Published: (2025)
by: Do, Heejin, et al.
Published: (2025)
MASEval: Extending Multi-Agent Evaluation from Models to Systems
by: Emde, Cornelius, et al.
Published: (2026)
by: Emde, Cornelius, et al.
Published: (2026)
Multimodal Cognitive Reframing Therapy via Multi-hop Psychotherapeutic Reasoning
by: Kim, Subin, et al.
Published: (2025)
by: Kim, Subin, et al.
Published: (2025)
C-SEO Bench: Does Conversational SEO Work?
by: Puerto, Haritz, et al.
Published: (2025)
by: Puerto, Haritz, et al.
Published: (2025)
Calibrating Large Language Models Using Their Generations Only
by: Ulmer, Dennis, et al.
Published: (2024)
by: Ulmer, Dennis, et al.
Published: (2024)
TRAP: Targeted Random Adversarial Prompt Honeypot for Black-Box Identification
by: Gubri, Martin, et al.
Published: (2024)
by: Gubri, Martin, et al.
Published: (2024)
Prompt Architecture Determines Reasoning Quality: A Variable Isolation Study on the Car Wash Problem
by: Jo, Heejin
Published: (2026)
by: Jo, Heejin
Published: (2026)
Leveraging What's Overfixed: Post-Correction via LLM Grammatical Error Overcorrection
by: Park, Taehee, et al.
Published: (2025)
by: Park, Taehee, et al.
Published: (2025)
Dissecting Logical Reasoning in LLMs: A Fine-Grained Evaluation and Supervision Study
by: Zhou, Yujun, et al.
Published: (2025)
by: Zhou, Yujun, et al.
Published: (2025)
What Makes Good Multilingual Reasoning? Disentangling Reasoning Traces with Measurable Features
by: Ki, Dayeon, et al.
Published: (2026)
by: Ki, Dayeon, et al.
Published: (2026)
Code-Switching Curriculum Learning for Multilingual Transfer in LLMs
by: Yoo, Haneul, et al.
Published: (2024)
by: Yoo, Haneul, et al.
Published: (2024)
Autoregressive Score Generation for Multi-trait Essay Scoring
by: Do, Heejin, et al.
Published: (2024)
by: Do, Heejin, et al.
Published: (2024)
RFEval: Benchmarking Reasoning Faithfulness under Counterfactual Reasoning Intervention in Large Reasoning Models
by: Han, Yunseok, et al.
Published: (2026)
by: Han, Yunseok, et al.
Published: (2026)
Autoregressive Multi-trait Essay Scoring via Reinforcement Learning with Scoring-aware Multiple Rewards
by: Do, Heejin, et al.
Published: (2024)
by: Do, Heejin, et al.
Published: (2024)
A Step Towards Mixture of Grader: Statistical Analysis of Existing Automatic Evaluation Metrics
by: Soh, Yun Joon, et al.
Published: (2024)
by: Soh, Yun Joon, et al.
Published: (2024)
LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models
by: Hao, Shibo, et al.
Published: (2024)
by: Hao, Shibo, et al.
Published: (2024)
MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs
by: Wu, Juncheng, et al.
Published: (2025)
by: Wu, Juncheng, et al.
Published: (2025)
Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter?
by: Tyagi, Nemika, et al.
Published: (2024)
by: Tyagi, Nemika, et al.
Published: (2024)
Large-Scale Aspect-Based Sentiment Analysis with Reasoning-Infused LLMs
by: Liskowski, Paweł, et al.
Published: (2026)
by: Liskowski, Paweł, et al.
Published: (2026)
From Building Blocks to Planning: Multi-Step Spatial Reasoning in LLMs with Reinforcement Learning
by: Tahmasbi, Amir, et al.
Published: (2025)
by: Tahmasbi, Amir, et al.
Published: (2025)
Do LLMs Really Think Step-by-step In Implicit Reasoning?
by: Yu, Yijiong
Published: (2024)
by: Yu, Yijiong
Published: (2024)
SuperCLUE-Math6: Graded Multi-Step Math Reasoning Benchmark for LLMs in Chinese
by: Xu, Liang, et al.
Published: (2024)
by: Xu, Liang, et al.
Published: (2024)
Tracking the Limits of Knowledge Propagation: How LLMs Fail at Multi-Step Reasoning with Conflicting Knowledge
by: Feng, Yiyang, et al.
Published: (2026)
by: Feng, Yiyang, et al.
Published: (2026)
MEME: Multi-entity & Evolving Memory Evaluation
by: Jung, Seokwon, et al.
Published: (2026)
by: Jung, Seokwon, et al.
Published: (2026)
Are Large Language Models Really Good Logical Reasoners? A Comprehensive Evaluation and Beyond
by: Xu, Fangzhi, et al.
Published: (2023)
by: Xu, Fangzhi, et al.
Published: (2023)
Reasoning or Not? A Comprehensive Evaluation of Reasoning LLMs for Dialogue Summarization
by: Jin, Keyan, et al.
Published: (2025)
by: Jin, Keyan, et al.
Published: (2025)
R3-RAG: Learning Step-by-Step Reasoning and Retrieval for LLMs via Reinforcement Learning
by: Li, Yuan, et al.
Published: (2025)
by: Li, Yuan, et al.
Published: (2025)
Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models
by: Patel, Nisarg, et al.
Published: (2024)
by: Patel, Nisarg, et al.
Published: (2024)
Dissecting Failure Dynamics in Large Language Model Reasoning
by: Zhu, Wei, et al.
Published: (2026)
by: Zhu, Wei, et al.
Published: (2026)
Dissecting Tool-Integrated Reasoning: An Empirical Study and Analysis
by: Zhao, Yufeng, et al.
Published: (2025)
by: Zhao, Yufeng, et al.
Published: (2025)
Are Large Reasoning Models Good Translation Evaluators? Analysis and Performance Boost
by: Zhan, Runzhe, et al.
Published: (2025)
by: Zhan, Runzhe, et al.
Published: (2025)
Multi-Dimensional Optimization for Text Summarization via Reinforcement Learning
by: Ryu, Sangwon, et al.
Published: (2024)
by: Ryu, Sangwon, et al.
Published: (2024)
LiteStage: Latency-aware Layer Skipping for Multi-stage Reasoning
by: Kang, Beomseok, et al.
Published: (2025)
by: Kang, Beomseok, et al.
Published: (2025)
Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models
by: Gu, Xiaojie, et al.
Published: (2026)
by: Gu, Xiaojie, et al.
Published: (2026)
ECG-Reasoning-Benchmark: A Benchmark for Evaluating Clinical Reasoning Capabilities in ECG Interpretation
by: Oh, Jungwoo, et al.
Published: (2026)
by: Oh, Jungwoo, et al.
Published: (2026)
Similar Items
-
Oops, Wait: Token-Level Signals as a Lens into LLM Reasoning
by: Hwang, Jaehui, et al.
Published: (2026) -
Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers
by: Green, Tommaso, et al.
Published: (2025) -
Dr.LLM: Dynamic Layer Routing in LLMs
by: Heakl, Ahmed, et al.
Published: (2025) -
Scaling Up Membership Inference: When and How Attacks Succeed on Large Language Models
by: Puerto, Haritz, et al.
Published: (2024) -
StoryCoder: Narrative Reformulation for Structured Reasoning in LLM Code Generation
by: Jang, Geonhui, et al.
Published: (2026)