Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yun, Taewon, Shin, Jisu, Choi, Jeonghwan, Bang, Seunghwan, Song, Hwanjun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning to Verify Summary Facts with Fine-Grained LLM Feedback
von: Oh, Jihwan, et al.
Veröffentlicht: (2024)
von: Oh, Jihwan, et al.
Veröffentlicht: (2024)
ReFeed: Multi-dimensional Summarization Refinement with Reflective Reasoning on Feedback
von: Yun, Taewon, et al.
Veröffentlicht: (2025)
von: Yun, Taewon, et al.
Veröffentlicht: (2025)
Aligning Extraction and Generation for Robust Retrieval-Augmented Generation
von: Song, Hwanjun, et al.
Veröffentlicht: (2025)
von: Song, Hwanjun, et al.
Veröffentlicht: (2025)
What Makes a Sale? Rethinking End-to-End Seller--Buyer Retail Dynamics with LLM Agents
von: Choi, Jeonghwan, et al.
Veröffentlicht: (2026)
von: Choi, Jeonghwan, et al.
Veröffentlicht: (2026)
From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
von: Deng, Yuntian, et al.
Veröffentlicht: (2024)
von: Deng, Yuntian, et al.
Veröffentlicht: (2024)
UniSumEval: Towards Unified, Fine-Grained, Multi-Dimensional Summarization Evaluation for LLMs
von: Lee, Yuho, et al.
Veröffentlicht: (2024)
von: Lee, Yuho, et al.
Veröffentlicht: (2024)
Deconstructing Long Chain-of-Thought: A Structured Reasoning Optimization Framework for Long CoT Distillation
von: Luo, Yijia, et al.
Veröffentlicht: (2025)
von: Luo, Yijia, et al.
Veröffentlicht: (2025)
Reasoning over Video: Evaluating How MLLMs Extract, Integrate, and Reconstruct Spatiotemporal Evidence
von: Bang, Seunghwan, et al.
Veröffentlicht: (2026)
von: Bang, Seunghwan, et al.
Veröffentlicht: (2026)
Improve Student's Reasoning Generalizability through Cascading Decomposed CoTs Distillation
von: Dai, Chengwei, et al.
Veröffentlicht: (2024)
von: Dai, Chengwei, et al.
Veröffentlicht: (2024)
The Quest for Efficient Reasoning: A Data-Centric Benchmark to CoT Distillation
von: Zhang, Ruichen, et al.
Veröffentlicht: (2025)
von: Zhang, Ruichen, et al.
Veröffentlicht: (2025)
LLM-based User Profile Management for Recommender System
von: Bang, Seunghwan, et al.
Veröffentlicht: (2025)
von: Bang, Seunghwan, et al.
Veröffentlicht: (2025)
"The Whole Is Greater Than the Sum of Its Parts": A Compatibility-Aware Multi-Teacher CoT Distillation Framework
von: Cui, Jin, et al.
Veröffentlicht: (2026)
von: Cui, Jin, et al.
Veröffentlicht: (2026)
Embodied CoT Distillation From LLM To Off-the-shelf Agents
von: Choi, Wonje, et al.
Veröffentlicht: (2024)
von: Choi, Wonje, et al.
Veröffentlicht: (2024)
Knowledge-Augmented Long-CoT Generation for Complex Biomolecular Reasoning
von: Lyu, Tianwen, et al.
Veröffentlicht: (2025)
von: Lyu, Tianwen, et al.
Veröffentlicht: (2025)
Efficient Long CoT Reasoning in Small Language Models
von: Wang, Zhaoyang, et al.
Veröffentlicht: (2025)
von: Wang, Zhaoyang, et al.
Veröffentlicht: (2025)
Alignment Tuning for Large Language Models: A Data-Centric Lens on Alignment Data Pipelines
von: Song, Hwanjun
Veröffentlicht: (2026)
von: Song, Hwanjun
Veröffentlicht: (2026)
STEPER: Step-wise Knowledge Distillation for Enhancing Reasoning Ability in Multi-Step Retrieval-Augmented Language Models
von: Lee, Kyumin, et al.
Veröffentlicht: (2025)
von: Lee, Kyumin, et al.
Veröffentlicht: (2025)
Learning to Summarize from LLM-generated Feedback
von: Song, Hwanjun, et al.
Veröffentlicht: (2024)
von: Song, Hwanjun, et al.
Veröffentlicht: (2024)
Completing Missing Annotation: Multi-Agent Debate for Accurate and Scalable Relevant Assessment for IR Benchmarks
von: Ban, Minjeong, et al.
Veröffentlicht: (2026)
von: Ban, Minjeong, et al.
Veröffentlicht: (2026)
Do Latent-CoT Models Think Step-by-Step? A Mechanistic Study on Sequential Reasoning Tasks
von: Liang, Jia, et al.
Veröffentlicht: (2026)
von: Liang, Jia, et al.
Veröffentlicht: (2026)
Towards a Holistic and Automated Evaluation Framework for Multi-Level Comprehension of LLMs in Book-Length Contexts
von: Deng, Jiaqi, et al.
Veröffentlicht: (2025)
von: Deng, Jiaqi, et al.
Veröffentlicht: (2025)
Hindsight Hint Distillation: Scaffolded Reasoning for SWE Agents from CoT-free Answers
von: Wang, Shengjie, et al.
Veröffentlicht: (2026)
von: Wang, Shengjie, et al.
Veröffentlicht: (2026)
Select2Reason: Efficient Instruction-Tuning Data Selection for Long-CoT Reasoning
von: Yang, Cehao, et al.
Veröffentlicht: (2025)
von: Yang, Cehao, et al.
Veröffentlicht: (2025)
Evaluating LLM Reasoning Beyond Correctness and CoT
von: Abbasloo, Soheil
Veröffentlicht: (2025)
von: Abbasloo, Soheil
Veröffentlicht: (2025)
CAP-CoT: Cycle Adversarial Prompt for Improving Chain of Thoughts in LLM Reasoning
von: Chen, Shuxu, et al.
Veröffentlicht: (2026)
von: Chen, Shuxu, et al.
Veröffentlicht: (2026)
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
von: Lai, Xin, et al.
Veröffentlicht: (2024)
von: Lai, Xin, et al.
Veröffentlicht: (2024)
Towards Multi-dimensional Evaluation of LLM Summarization across Domains and Languages
von: Min, Hyangsuk, et al.
Veröffentlicht: (2025)
von: Min, Hyangsuk, et al.
Veröffentlicht: (2025)
CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning
von: Dong, Qihua, et al.
Veröffentlicht: (2025)
von: Dong, Qihua, et al.
Veröffentlicht: (2025)
SOD: Step-wise On-policy Distillation for Small Language Model Agents
von: Zhong, Qiyong, et al.
Veröffentlicht: (2026)
von: Zhong, Qiyong, et al.
Veröffentlicht: (2026)
Hybrid Distillation with CoT Guidance for Edge-Drone Control Code Generation
von: Feng, Yizhan, et al.
Veröffentlicht: (2026)
von: Feng, Yizhan, et al.
Veröffentlicht: (2026)
Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT
von: Asfour, Alaa, et al.
Veröffentlicht: (2026)
von: Asfour, Alaa, et al.
Veröffentlicht: (2026)
LegalReasoner: Step-wised Verification-Correction for Legal Judgment Reasoning
von: Shi, Weijie, et al.
Veröffentlicht: (2025)
von: Shi, Weijie, et al.
Veröffentlicht: (2025)
The Potential of CoT for Reasoning: A Closer Look at Trace Dynamics
von: Bachmann, Gregor, et al.
Veröffentlicht: (2026)
von: Bachmann, Gregor, et al.
Veröffentlicht: (2026)
SLAT: Segment-Level Adaptive Trimming for Efficient CoT Reasoning
von: Yao, Jian, et al.
Veröffentlicht: (2026)
von: Yao, Jian, et al.
Veröffentlicht: (2026)
Reasoning as an Attack Surface: Adaptive Evolutionary CoT Jailbreaks for LLMs
von: Li, Jianan, et al.
Veröffentlicht: (2026)
von: Li, Jianan, et al.
Veröffentlicht: (2026)
Reconsidering Overthinking: Penalizing Internal and External Redundancy in CoT Reasoning
von: Hong, Jialiang, et al.
Veröffentlicht: (2025)
von: Hong, Jialiang, et al.
Veröffentlicht: (2025)
D-SCoRE: Document-Centric Segmentation and CoT Reasoning with Structured Export for QA-CoT Data Generation
von: Zhou, Weibo, et al.
Veröffentlicht: (2025)
von: Zhou, Weibo, et al.
Veröffentlicht: (2025)
Teacher as a Lenient Expert: Teacher-Agnostic Data-Free Knowledge Distillation
von: Shin, Hyunjune, et al.
Veröffentlicht: (2024)
von: Shin, Hyunjune, et al.
Veröffentlicht: (2024)
VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought
von: Lim, Byeonggeuk, et al.
Veröffentlicht: (2026)
von: Lim, Byeonggeuk, et al.
Veröffentlicht: (2026)
Diagnosing Harmful Continuation in Answer-Correct Long-CoT Training Traces
von: He, Chen, et al.
Veröffentlicht: (2026)
von: He, Chen, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Learning to Verify Summary Facts with Fine-Grained LLM Feedback
von: Oh, Jihwan, et al.
Veröffentlicht: (2024) -
ReFeed: Multi-dimensional Summarization Refinement with Reflective Reasoning on Feedback
von: Yun, Taewon, et al.
Veröffentlicht: (2025) -
Aligning Extraction and Generation for Robust Retrieval-Augmented Generation
von: Song, Hwanjun, et al.
Veröffentlicht: (2025) -
What Makes a Sale? Rethinking End-to-End Seller--Buyer Retail Dynamics with LLM Agents
von: Choi, Jeonghwan, et al.
Veröffentlicht: (2026) -
From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
von: Deng, Yuntian, et al.
Veröffentlicht: (2024)