Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding
Fuente:
arXiv
Salvato in:
| Autori principali: | Yun, Taewon, Shin, Jisu, Choi, Jeonghwan, Bang, Seunghwan, Song, Hwanjun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Learning to Verify Summary Facts with Fine-Grained LLM Feedback
di: Oh, Jihwan, et al.
Pubblicazione: (2024)
di: Oh, Jihwan, et al.
Pubblicazione: (2024)
ReFeed: Multi-dimensional Summarization Refinement with Reflective Reasoning on Feedback
di: Yun, Taewon, et al.
Pubblicazione: (2025)
di: Yun, Taewon, et al.
Pubblicazione: (2025)
Aligning Extraction and Generation for Robust Retrieval-Augmented Generation
di: Song, Hwanjun, et al.
Pubblicazione: (2025)
di: Song, Hwanjun, et al.
Pubblicazione: (2025)
What Makes a Sale? Rethinking End-to-End Seller--Buyer Retail Dynamics with LLM Agents
di: Choi, Jeonghwan, et al.
Pubblicazione: (2026)
di: Choi, Jeonghwan, et al.
Pubblicazione: (2026)
From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
di: Deng, Yuntian, et al.
Pubblicazione: (2024)
di: Deng, Yuntian, et al.
Pubblicazione: (2024)
UniSumEval: Towards Unified, Fine-Grained, Multi-Dimensional Summarization Evaluation for LLMs
di: Lee, Yuho, et al.
Pubblicazione: (2024)
di: Lee, Yuho, et al.
Pubblicazione: (2024)
Deconstructing Long Chain-of-Thought: A Structured Reasoning Optimization Framework for Long CoT Distillation
di: Luo, Yijia, et al.
Pubblicazione: (2025)
di: Luo, Yijia, et al.
Pubblicazione: (2025)
Reasoning over Video: Evaluating How MLLMs Extract, Integrate, and Reconstruct Spatiotemporal Evidence
di: Bang, Seunghwan, et al.
Pubblicazione: (2026)
di: Bang, Seunghwan, et al.
Pubblicazione: (2026)
Improve Student's Reasoning Generalizability through Cascading Decomposed CoTs Distillation
di: Dai, Chengwei, et al.
Pubblicazione: (2024)
di: Dai, Chengwei, et al.
Pubblicazione: (2024)
The Quest for Efficient Reasoning: A Data-Centric Benchmark to CoT Distillation
di: Zhang, Ruichen, et al.
Pubblicazione: (2025)
di: Zhang, Ruichen, et al.
Pubblicazione: (2025)
LLM-based User Profile Management for Recommender System
di: Bang, Seunghwan, et al.
Pubblicazione: (2025)
di: Bang, Seunghwan, et al.
Pubblicazione: (2025)
"The Whole Is Greater Than the Sum of Its Parts": A Compatibility-Aware Multi-Teacher CoT Distillation Framework
di: Cui, Jin, et al.
Pubblicazione: (2026)
di: Cui, Jin, et al.
Pubblicazione: (2026)
Embodied CoT Distillation From LLM To Off-the-shelf Agents
di: Choi, Wonje, et al.
Pubblicazione: (2024)
di: Choi, Wonje, et al.
Pubblicazione: (2024)
Knowledge-Augmented Long-CoT Generation for Complex Biomolecular Reasoning
di: Lyu, Tianwen, et al.
Pubblicazione: (2025)
di: Lyu, Tianwen, et al.
Pubblicazione: (2025)
Efficient Long CoT Reasoning in Small Language Models
di: Wang, Zhaoyang, et al.
Pubblicazione: (2025)
di: Wang, Zhaoyang, et al.
Pubblicazione: (2025)
Alignment Tuning for Large Language Models: A Data-Centric Lens on Alignment Data Pipelines
di: Song, Hwanjun
Pubblicazione: (2026)
di: Song, Hwanjun
Pubblicazione: (2026)
STEPER: Step-wise Knowledge Distillation for Enhancing Reasoning Ability in Multi-Step Retrieval-Augmented Language Models
di: Lee, Kyumin, et al.
Pubblicazione: (2025)
di: Lee, Kyumin, et al.
Pubblicazione: (2025)
Learning to Summarize from LLM-generated Feedback
di: Song, Hwanjun, et al.
Pubblicazione: (2024)
di: Song, Hwanjun, et al.
Pubblicazione: (2024)
Completing Missing Annotation: Multi-Agent Debate for Accurate and Scalable Relevant Assessment for IR Benchmarks
di: Ban, Minjeong, et al.
Pubblicazione: (2026)
di: Ban, Minjeong, et al.
Pubblicazione: (2026)
Do Latent-CoT Models Think Step-by-Step? A Mechanistic Study on Sequential Reasoning Tasks
di: Liang, Jia, et al.
Pubblicazione: (2026)
di: Liang, Jia, et al.
Pubblicazione: (2026)
Towards a Holistic and Automated Evaluation Framework for Multi-Level Comprehension of LLMs in Book-Length Contexts
di: Deng, Jiaqi, et al.
Pubblicazione: (2025)
di: Deng, Jiaqi, et al.
Pubblicazione: (2025)
Hindsight Hint Distillation: Scaffolded Reasoning for SWE Agents from CoT-free Answers
di: Wang, Shengjie, et al.
Pubblicazione: (2026)
di: Wang, Shengjie, et al.
Pubblicazione: (2026)
Select2Reason: Efficient Instruction-Tuning Data Selection for Long-CoT Reasoning
di: Yang, Cehao, et al.
Pubblicazione: (2025)
di: Yang, Cehao, et al.
Pubblicazione: (2025)
Evaluating LLM Reasoning Beyond Correctness and CoT
di: Abbasloo, Soheil
Pubblicazione: (2025)
di: Abbasloo, Soheil
Pubblicazione: (2025)
CAP-CoT: Cycle Adversarial Prompt for Improving Chain of Thoughts in LLM Reasoning
di: Chen, Shuxu, et al.
Pubblicazione: (2026)
di: Chen, Shuxu, et al.
Pubblicazione: (2026)
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
di: Lai, Xin, et al.
Pubblicazione: (2024)
di: Lai, Xin, et al.
Pubblicazione: (2024)
Towards Multi-dimensional Evaluation of LLM Summarization across Domains and Languages
di: Min, Hyangsuk, et al.
Pubblicazione: (2025)
di: Min, Hyangsuk, et al.
Pubblicazione: (2025)
CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning
di: Dong, Qihua, et al.
Pubblicazione: (2025)
di: Dong, Qihua, et al.
Pubblicazione: (2025)
SOD: Step-wise On-policy Distillation for Small Language Model Agents
di: Zhong, Qiyong, et al.
Pubblicazione: (2026)
di: Zhong, Qiyong, et al.
Pubblicazione: (2026)
Hybrid Distillation with CoT Guidance for Edge-Drone Control Code Generation
di: Feng, Yizhan, et al.
Pubblicazione: (2026)
di: Feng, Yizhan, et al.
Pubblicazione: (2026)
Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT
di: Asfour, Alaa, et al.
Pubblicazione: (2026)
di: Asfour, Alaa, et al.
Pubblicazione: (2026)
LegalReasoner: Step-wised Verification-Correction for Legal Judgment Reasoning
di: Shi, Weijie, et al.
Pubblicazione: (2025)
di: Shi, Weijie, et al.
Pubblicazione: (2025)
The Potential of CoT for Reasoning: A Closer Look at Trace Dynamics
di: Bachmann, Gregor, et al.
Pubblicazione: (2026)
di: Bachmann, Gregor, et al.
Pubblicazione: (2026)
SLAT: Segment-Level Adaptive Trimming for Efficient CoT Reasoning
di: Yao, Jian, et al.
Pubblicazione: (2026)
di: Yao, Jian, et al.
Pubblicazione: (2026)
Reasoning as an Attack Surface: Adaptive Evolutionary CoT Jailbreaks for LLMs
di: Li, Jianan, et al.
Pubblicazione: (2026)
di: Li, Jianan, et al.
Pubblicazione: (2026)
Reconsidering Overthinking: Penalizing Internal and External Redundancy in CoT Reasoning
di: Hong, Jialiang, et al.
Pubblicazione: (2025)
di: Hong, Jialiang, et al.
Pubblicazione: (2025)
D-SCoRE: Document-Centric Segmentation and CoT Reasoning with Structured Export for QA-CoT Data Generation
di: Zhou, Weibo, et al.
Pubblicazione: (2025)
di: Zhou, Weibo, et al.
Pubblicazione: (2025)
Teacher as a Lenient Expert: Teacher-Agnostic Data-Free Knowledge Distillation
di: Shin, Hyunjune, et al.
Pubblicazione: (2024)
di: Shin, Hyunjune, et al.
Pubblicazione: (2024)
VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought
di: Lim, Byeonggeuk, et al.
Pubblicazione: (2026)
di: Lim, Byeonggeuk, et al.
Pubblicazione: (2026)
Diagnosing Harmful Continuation in Answer-Correct Long-CoT Training Traces
di: He, Chen, et al.
Pubblicazione: (2026)
di: He, Chen, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Learning to Verify Summary Facts with Fine-Grained LLM Feedback
di: Oh, Jihwan, et al.
Pubblicazione: (2024) -
ReFeed: Multi-dimensional Summarization Refinement with Reflective Reasoning on Feedback
di: Yun, Taewon, et al.
Pubblicazione: (2025) -
Aligning Extraction and Generation for Robust Retrieval-Augmented Generation
di: Song, Hwanjun, et al.
Pubblicazione: (2025) -
What Makes a Sale? Rethinking End-to-End Seller--Buyer Retail Dynamics with LLM Agents
di: Choi, Jeonghwan, et al.
Pubblicazione: (2026) -
From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
di: Deng, Yuntian, et al.
Pubblicazione: (2024)