Hindsight Hint Distillation: Scaffolded Reasoning for SWE Agents from CoT-free Answers
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Shengjie, Li, Guanghe, Yang, Zonghan, Gao, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HEAL: Hindsight Entropy-Assisted Learning for Reasoning Distillation
by: Zhang, Wenjing, et al.
Published: (2026)
by: Zhang, Wenjing, et al.
Published: (2026)
The Quest for Efficient Reasoning: A Data-Centric Benchmark to CoT Distillation
by: Zhang, Ruichen, et al.
Published: (2025)
by: Zhang, Ruichen, et al.
Published: (2025)
How Likely Do LLMs with CoT Mimic Human Reasoning?
by: Bao, Guangsheng, et al.
Published: (2024)
by: Bao, Guangsheng, et al.
Published: (2024)
Hindsight Credit Assignment for Long-Horizon LLM Agents
by: Tan, Hui-Ze, et al.
Published: (2026)
by: Tan, Hui-Ze, et al.
Published: (2026)
Shorter Thoughts, Same Answers: Difficulty-Scaled Segment-Wise RL for CoT Compression
by: Tian, Ye, et al.
Published: (2026)
by: Tian, Ye, et al.
Published: (2026)
HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agents
by: Yeo, Woongyeng, et al.
Published: (2026)
by: Yeo, Woongyeng, et al.
Published: (2026)
Exploring the Limitations of Mamba in COPY and CoT Reasoning
by: Ren, Ruifeng, et al.
Published: (2024)
by: Ren, Ruifeng, et al.
Published: (2024)
Hindsight Preference Learning for Offline Preference-based Reinforcement Learning
by: Gao, Chen-Xiao, et al.
Published: (2024)
by: Gao, Chen-Xiao, et al.
Published: (2024)
From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
by: Deng, Yuntian, et al.
Published: (2024)
by: Deng, Yuntian, et al.
Published: (2024)
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason
by: Zhang, Kaiyi, et al.
Published: (2025)
by: Zhang, Kaiyi, et al.
Published: (2025)
The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation
by: Lan, Yifan, et al.
Published: (2026)
by: Lan, Yifan, et al.
Published: (2026)
GCHR : Goal-Conditioned Hindsight Regularization for Sample-Efficient Reinforcement Learning
by: Lei, Xing, et al.
Published: (2025)
by: Lei, Xing, et al.
Published: (2025)
The Ends Justify the Thoughts: RL-Induced Motivated Reasoning in LLM CoTs
by: Howe, Nikolaus, et al.
Published: (2025)
by: Howe, Nikolaus, et al.
Published: (2025)
LatentChem: From Textual CoT to Latent Thinking in Chemical Reasoning
by: Ye, Xinwu, et al.
Published: (2026)
by: Ye, Xinwu, et al.
Published: (2026)
EasyInsert: A Data-Efficient and Generalizable Insertion Policy
by: Li, Guanghe, et al.
Published: (2025)
by: Li, Guanghe, et al.
Published: (2025)
To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
by: Sprague, Zayne, et al.
Published: (2024)
by: Sprague, Zayne, et al.
Published: (2024)
RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?
by: Xu, Haotian, et al.
Published: (2025)
by: Xu, Haotian, et al.
Published: (2025)
Hindsight Preference Optimization for Financial Time Series Advisory
by: Cui, Yanwei, et al.
Published: (2026)
by: Cui, Yanwei, et al.
Published: (2026)
Cross Domain Evaluation of Multimodal Chain-of-Thought Reasoning of different datasets into the Amazon CoT Framework
by: Tiwari, Nitya, et al.
Published: (2025)
by: Tiwari, Nitya, et al.
Published: (2025)
ReAct Meets ActRe: When Language Agents Enjoy Training Data Autonomy
by: Yang, Zonghan, et al.
Published: (2024)
by: Yang, Zonghan, et al.
Published: (2024)
Do Latent-CoT Models Think Step-by-Step? A Mechanistic Study on Sequential Reasoning Tasks
by: Liang, Jia, et al.
Published: (2026)
by: Liang, Jia, et al.
Published: (2026)
Data Shifts Hurt CoT: A Theoretical Study
by: Yin, Lang, et al.
Published: (2025)
by: Yin, Lang, et al.
Published: (2025)
Interactive Dialogue Agents via Reinforcement Learning on Hindsight Regenerations
by: Hong, Joey, et al.
Published: (2024)
by: Hong, Joey, et al.
Published: (2024)
Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents
by: Huang, Jiawei, et al.
Published: (2026)
by: Huang, Jiawei, et al.
Published: (2026)
DiffStitch: Boosting Offline Reinforcement Learning with Diffusion-based Trajectory Stitching
by: Li, Guanghe, et al.
Published: (2024)
by: Li, Guanghe, et al.
Published: (2024)
CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring
by: Arnav, Benjamin, et al.
Published: (2025)
by: Arnav, Benjamin, et al.
Published: (2025)
Beyond Answers: Transferring Reasoning Capabilities to Smaller LLMs Using Multi-Teacher Knowledge Distillation
by: Tian, Yijun, et al.
Published: (2024)
by: Tian, Yijun, et al.
Published: (2024)
Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?
by: Xia, Chunqiu Steven, et al.
Published: (2025)
by: Xia, Chunqiu Steven, et al.
Published: (2025)
Compositional Generalization from Learned Skills via CoT Training: A Theoretical and Structural Analysis for Reasoning
by: Yao, Xinhao, et al.
Published: (2025)
by: Yao, Xinhao, et al.
Published: (2025)
SWE-Bench-CL: Continual Learning for Coding Agents
by: Joshi, Thomas, et al.
Published: (2025)
by: Joshi, Thomas, et al.
Published: (2025)
UML-CoT: Structured Reasoning and Planning with Unified Modeling Language for Robotic Room Cleaning
by: Chen, Hongyu, et al.
Published: (2025)
by: Chen, Hongyu, et al.
Published: (2025)
Beyond the Answer: Decoding the Behavior of LLMs as Scientific Reasoners
by: Pandey, Rohan, et al.
Published: (2026)
by: Pandey, Rohan, et al.
Published: (2026)
Mitigating Distribution Sharpening in Math RLVR via Distribution-Aligned Hint Synthesis and Backward Hint Annealing
by: Xie, Pei-Xi, et al.
Published: (2026)
by: Xie, Pei-Xi, et al.
Published: (2026)
What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents
by: Li, Xiaozhe, et al.
Published: (2026)
by: Li, Xiaozhe, et al.
Published: (2026)
Plantain: Plan-Answer Interleaved Reasoning
by: Liang, Anthony, et al.
Published: (2025)
by: Liang, Anthony, et al.
Published: (2025)
Sample-Efficient Online Learning in LM Agents via Hindsight Trajectory Rewriting
by: Hu, Michael Y., et al.
Published: (2025)
by: Hu, Michael Y., et al.
Published: (2025)
Edge-free but Structure-aware: Prototype-Guided Knowledge Distillation from GNNs to MLPs
by: Wu, Taiqiang, et al.
Published: (2023)
by: Wu, Taiqiang, et al.
Published: (2023)
Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning
by: Zhou, Yang, et al.
Published: (2025)
by: Zhou, Yang, et al.
Published: (2025)
PepThink-R1: LLM for Interpretable Cyclic Peptide Optimization with CoT SFT and Reinforcement Learning
by: Wang, Ruheng, et al.
Published: (2025)
by: Wang, Ruheng, et al.
Published: (2025)
Hindsight is 20/20: Building Agent Memory that Retains, Recalls, and Reflects
by: Latimer, Chris, et al.
Published: (2025)
by: Latimer, Chris, et al.
Published: (2025)
Similar Items
-
HEAL: Hindsight Entropy-Assisted Learning for Reasoning Distillation
by: Zhang, Wenjing, et al.
Published: (2026) -
The Quest for Efficient Reasoning: A Data-Centric Benchmark to CoT Distillation
by: Zhang, Ruichen, et al.
Published: (2025) -
How Likely Do LLMs with CoT Mimic Human Reasoning?
by: Bao, Guangsheng, et al.
Published: (2024) -
Hindsight Credit Assignment for Long-Horizon LLM Agents
by: Tan, Hui-Ze, et al.
Published: (2026) -
Shorter Thoughts, Same Answers: Difficulty-Scaled Segment-Wise RL for CoT Compression
by: Tian, Ye, et al.
Published: (2026)