Evaluating Step-by-step Reasoning Traces: A Survey
Fuente:
arXiv
Salvato in:
| Autori principali: | Lee, Jinu, Hockenmaier, Julia |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ReasoningFlow: Semantic Structure of Complex Reasoning Traces
di: Lee, Jinu, et al.
Pubblicazione: (2025)
di: Lee, Jinu, et al.
Pubblicazione: (2025)
Evaluating Legal Reasoning Traces with Legal Issue Tree Rubrics
di: Lee, Jinu, et al.
Pubblicazione: (2025)
di: Lee, Jinu, et al.
Pubblicazione: (2025)
Scaling Evaluation-time Compute with Reasoning Models as Evaluators
di: Kim, Seungone, et al.
Pubblicazione: (2025)
di: Kim, Seungone, et al.
Pubblicazione: (2025)
Grid Spatial Understanding: A Dataset for Textual Spatial Reasoning over Grids, Embodied Settings, and Coordinate Structures
di: Sidhu, Risham, et al.
Pubblicazione: (2026)
di: Sidhu, Risham, et al.
Pubblicazione: (2026)
Entailment-Preserving First-order Logic Representations in Natural Language Entailment
di: Lee, Jinu, et al.
Pubblicazione: (2025)
di: Lee, Jinu, et al.
Pubblicazione: (2025)
Rating Roulette: Self-Inconsistency in LLM-As-A-Judge Frameworks
di: Haldar, Rajarshi, et al.
Pubblicazione: (2025)
di: Haldar, Rajarshi, et al.
Pubblicazione: (2025)
SymBa: Symbolic Backward Chaining for Structured Natural Language Reasoning
di: Lee, Jinu, et al.
Pubblicazione: (2024)
di: Lee, Jinu, et al.
Pubblicazione: (2024)
Analyzing the Performance of Large Language Models on Code Summarization
di: Haldar, Rajarshi, et al.
Pubblicazione: (2024)
di: Haldar, Rajarshi, et al.
Pubblicazione: (2024)
A Multi-Perspective Architecture for Semantic Code Search
di: Haldar, Rajarshi, et al.
Pubblicazione: (2020)
di: Haldar, Rajarshi, et al.
Pubblicazione: (2020)
Do LLMs Really Think Step-by-step In Implicit Reasoning?
di: Yu, Yijiong
Pubblicazione: (2024)
di: Yu, Yijiong
Pubblicazione: (2024)
Evaluating Step-by-Step Reasoning through Symbolic Verification
di: Zhang, Yi-Fan, et al.
Pubblicazione: (2022)
di: Zhang, Yi-Fan, et al.
Pubblicazione: (2022)
How Reliable are Causal Probing Interventions?
di: Canby, Marc, et al.
Pubblicazione: (2024)
di: Canby, Marc, et al.
Pubblicazione: (2024)
RAISE: Enhancing Scientific Reasoning in LLMs via Step-by-Step Retrieval
di: Oh, Minhae, et al.
Pubblicazione: (2025)
di: Oh, Minhae, et al.
Pubblicazione: (2025)
Large Language Model for Discrete Optimization Problems: Evaluation and Step-by-step Reasoning
di: Qian, Tianhao, et al.
Pubblicazione: (2026)
di: Qian, Tianhao, et al.
Pubblicazione: (2026)
ConCISE: Confidence-guided Compression in Step-by-step Efficient Reasoning
di: Qiao, Ziqing, et al.
Pubblicazione: (2025)
di: Qiao, Ziqing, et al.
Pubblicazione: (2025)
LegalSearchLM: Rethinking Legal Case Retrieval as Legal Elements Generation
di: Kim, Chaeeun, et al.
Pubblicazione: (2025)
di: Kim, Chaeeun, et al.
Pubblicazione: (2025)
Step-Tagging: Toward controlling the generation of Language Reasoning Models through step monitoring
di: Belkhiter, Yannis, et al.
Pubblicazione: (2025)
di: Belkhiter, Yannis, et al.
Pubblicazione: (2025)
StepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement Learning
di: Wang, Hao, et al.
Pubblicazione: (2026)
di: Wang, Hao, et al.
Pubblicazione: (2026)
LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models
di: Hao, Shibo, et al.
Pubblicazione: (2024)
di: Hao, Shibo, et al.
Pubblicazione: (2024)
MalruleLib: Large-Scale Executable Misconception Reasoning with Step Traces for Modeling Student Thinking in Mathematics
di: Chen, Xinghe, et al.
Pubblicazione: (2026)
di: Chen, Xinghe, et al.
Pubblicazione: (2026)
A Survey on Feedback-based Multi-step Reasoning for Large Language Models on Mathematics
di: Wei, Ting-Ruen, et al.
Pubblicazione: (2025)
di: Wei, Ting-Ruen, et al.
Pubblicazione: (2025)
ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question Answering
di: Molfese, Francesco Maria, et al.
Pubblicazione: (2025)
di: Molfese, Francesco Maria, et al.
Pubblicazione: (2025)
RAG-RL: Advancing Retrieval-Augmented Generation via RL and Curriculum Learning
di: Huang, Jerry, et al.
Pubblicazione: (2025)
di: Huang, Jerry, et al.
Pubblicazione: (2025)
ReasonOps: Operator Segmentation for LLM Reasoning Traces
di: Lee, Daniel, et al.
Pubblicazione: (2026)
di: Lee, Daniel, et al.
Pubblicazione: (2026)
In Their Own Words: Reasoning Traces Tailored for Small Models Make Them Better Reasoners
di: Kim, Jaehoon, et al.
Pubblicazione: (2025)
di: Kim, Jaehoon, et al.
Pubblicazione: (2025)
Reveal-Bangla: A Dataset for Cross-Lingual Multi-Step Reasoning Evaluation
di: Islam, Khondoker Ittehadul, et al.
Pubblicazione: (2025)
di: Islam, Khondoker Ittehadul, et al.
Pubblicazione: (2025)
A Survey of Explainable Knowledge Tracing
di: Bai, Yanhong, et al.
Pubblicazione: (2024)
di: Bai, Yanhong, et al.
Pubblicazione: (2024)
Mentor-KD: Making Small Language Models Better Multi-step Reasoners
di: Lee, Hojae, et al.
Pubblicazione: (2024)
di: Lee, Hojae, et al.
Pubblicazione: (2024)
Evaluating LLMs' Reasoning Over Ordered Procedural Steps
di: Anika, Adrita, et al.
Pubblicazione: (2025)
di: Anika, Adrita, et al.
Pubblicazione: (2025)
Investigating the Shortcomings of LLMs in Step-by-Step Legal Reasoning
di: Mishra, Venkatesh, et al.
Pubblicazione: (2025)
di: Mishra, Venkatesh, et al.
Pubblicazione: (2025)
RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code
di: Gautam, Dhruv, et al.
Pubblicazione: (2025)
di: Gautam, Dhruv, et al.
Pubblicazione: (2025)
Multi-Step Reasoning with Large Language Models, a Survey
di: Plaat, Aske, et al.
Pubblicazione: (2024)
di: Plaat, Aske, et al.
Pubblicazione: (2024)
Mind the Gap Between Spatial Reasoning and Acting! Step-by-Step Evaluation of Agents With Spatial-Gym
di: Kaesberg, Lars Benedikt, et al.
Pubblicazione: (2026)
di: Kaesberg, Lars Benedikt, et al.
Pubblicazione: (2026)
Is Sarcasm Detection A Step-by-Step Reasoning Process in Large Language Models?
di: Yao, Ben, et al.
Pubblicazione: (2024)
di: Yao, Ben, et al.
Pubblicazione: (2024)
Opening the Black Box: A Survey on the Mechanisms of Multi-Step Reasoning in Large Language Models
di: Pan, Liangming, et al.
Pubblicazione: (2026)
di: Pan, Liangming, et al.
Pubblicazione: (2026)
What Defines Good Reasoning in LLMs? Dissecting Reasoning Steps with Multi-Aspect Evaluation
di: Do, Heejin, et al.
Pubblicazione: (2025)
di: Do, Heejin, et al.
Pubblicazione: (2025)
Filtered Reasoning Score: Evaluating Reasoning Quality on a Model's Most-Confident Traces
di: Pathak, Manas, et al.
Pubblicazione: (2026)
di: Pathak, Manas, et al.
Pubblicazione: (2026)
Instruct Large Language Models to Generate Scientific Literature Survey Step by Step
di: Lai, Yuxuan, et al.
Pubblicazione: (2024)
di: Lai, Yuxuan, et al.
Pubblicazione: (2024)
Step Guided Reasoning: Improving Mathematical Reasoning using Guidance Generation and Step Reasoning
di: Cao, Lang, et al.
Pubblicazione: (2024)
di: Cao, Lang, et al.
Pubblicazione: (2024)
STEPER: Step-wise Knowledge Distillation for Enhancing Reasoning Ability in Multi-Step Retrieval-Augmented Language Models
di: Lee, Kyumin, et al.
Pubblicazione: (2025)
di: Lee, Kyumin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
ReasoningFlow: Semantic Structure of Complex Reasoning Traces
di: Lee, Jinu, et al.
Pubblicazione: (2025) -
Evaluating Legal Reasoning Traces with Legal Issue Tree Rubrics
di: Lee, Jinu, et al.
Pubblicazione: (2025) -
Scaling Evaluation-time Compute with Reasoning Models as Evaluators
di: Kim, Seungone, et al.
Pubblicazione: (2025) -
Grid Spatial Understanding: A Dataset for Textual Spatial Reasoning over Grids, Embodied Settings, and Coordinate Structures
di: Sidhu, Risham, et al.
Pubblicazione: (2026) -
Entailment-Preserving First-order Logic Representations in Natural Language Entailment
di: Lee, Jinu, et al.
Pubblicazione: (2025)