Distribution-Aligned Sequence Distillation for Superior Long-CoT Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yan, Shaotian, Liu, Kaiyuan, Shen, Chen, Wang, Bing, Fan, Sinan, Zhang, Jun, Wu, Yue, Wang, Zheng, Ye, Jieping |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Where Did This Sentence Come From? Tracing Provenance in LLM Reasoning Distillation
von: Liu, Kaiyuan, et al.
Veröffentlicht: (2025)
von: Liu, Kaiyuan, et al.
Veröffentlicht: (2025)
On the Step Length Confounding in LLM Reasoning Data Selection
von: Wang, Bing, et al.
Veröffentlicht: (2026)
von: Wang, Bing, et al.
Veröffentlicht: (2026)
Backtracking When It Strays: Mitigating Dual Exposure Biases in LLM Reasoning Distillation
von: Wang, Bing, et al.
Veröffentlicht: (2026)
von: Wang, Bing, et al.
Veröffentlicht: (2026)
Enhancing Chain-of-Thought Reasoning with Critical Representation Fine-tuning
von: Huang, Chenxi, et al.
Veröffentlicht: (2025)
von: Huang, Chenxi, et al.
Veröffentlicht: (2025)
Don't Take Things Out of Context: Attention Intervention for Enhancing Chain-of-Thought Reasoning in Large Language Models
von: Yan, Shaotian, et al.
Veröffentlicht: (2025)
von: Yan, Shaotian, et al.
Veröffentlicht: (2025)
Concise and Organized Perception Facilitates Reasoning in Large Language Models
von: Liu, Junjie, et al.
Veröffentlicht: (2023)
von: Liu, Junjie, et al.
Veröffentlicht: (2023)
SalaMAnder: Shapley-based Mathematical Expression Attribution and Metric for Chain-of-Thought Reasoning
von: Xin, Yue, et al.
Veröffentlicht: (2025)
von: Xin, Yue, et al.
Veröffentlicht: (2025)
Are Rationales Necessary and Sufficient? Tuning LLMs for Explainable Misinformation Detection
von: Wang, Bing, et al.
Veröffentlicht: (2026)
von: Wang, Bing, et al.
Veröffentlicht: (2026)
Prefix Teach, Suffix Fade: Local Teachability Collapse in Strong-to-Weak On-Policy Distillation
von: Liu, Kaiyuan, et al.
Veröffentlicht: (2026)
von: Liu, Kaiyuan, et al.
Veröffentlicht: (2026)
Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs
von: Pan, Zhiyu, et al.
Veröffentlicht: (2026)
von: Pan, Zhiyu, et al.
Veröffentlicht: (2026)
Beyond Isolated Capabilities: Bridging Long CoT Reasoning and Long-Context Understanding
von: Wang, Yifei
Veröffentlicht: (2025)
von: Wang, Yifei
Veröffentlicht: (2025)
CoT-Evo: Evolutionary Distillation of Chain-of-Thought for Scientific Reasoning
von: Feng, Kehua, et al.
Veröffentlicht: (2025)
von: Feng, Kehua, et al.
Veröffentlicht: (2025)
Efficient Long CoT Reasoning in Small Language Models
von: Wang, Zhaoyang, et al.
Veröffentlicht: (2025)
von: Wang, Zhaoyang, et al.
Veröffentlicht: (2025)
Can Pruning Improve Reasoning? Revisiting Long-CoT Compression with Capability in Mind for Better Reasoning
von: Zhao, Shangziqi, et al.
Veröffentlicht: (2025)
von: Zhao, Shangziqi, et al.
Veröffentlicht: (2025)
Investigating CoT Monitorability in Large Reasoning Models
von: Yang, Shu, et al.
Veröffentlicht: (2025)
von: Yang, Shu, et al.
Veröffentlicht: (2025)
Instance-adaptive Zero-shot Chain-of-Thought Prompting
von: Yuan, Xiaosong, et al.
Veröffentlicht: (2024)
von: Yuan, Xiaosong, et al.
Veröffentlicht: (2024)
Mol-R1: Towards Explicit Long-CoT Reasoning in Molecule Discovery
von: Li, Jiatong, et al.
Veröffentlicht: (2025)
von: Li, Jiatong, et al.
Veröffentlicht: (2025)
From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks
von: Zhang, Xiaofeng, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaofeng, et al.
Veröffentlicht: (2024)
CoT Vectors: Transferring and Probing the Reasoning Mechanisms of LLMs
von: Li, Li, et al.
Veröffentlicht: (2025)
von: Li, Li, et al.
Veröffentlicht: (2025)
Investigating Mysteries of CoT-Augmented Distillation
von: Wadhwa, Somin, et al.
Veröffentlicht: (2024)
von: Wadhwa, Somin, et al.
Veröffentlicht: (2024)
Robust Reasoning via Dynamic Token Selection for Distribution-Aligned Self-Distillation
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2026)
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2026)
Parrot: A Training Pipeline Enhances Both Program CoT and Natural Language CoT for Reasoning
von: Jin, Senjie, et al.
Veröffentlicht: (2025)
von: Jin, Senjie, et al.
Veröffentlicht: (2025)
Reasoning with OmniThought: A Large CoT Dataset with Verbosity and Cognitive Difficulty Annotations
von: Cai, Wenrui, et al.
Veröffentlicht: (2025)
von: Cai, Wenrui, et al.
Veröffentlicht: (2025)
Efficient Reasoning Through Suppression of Self-Affirmation Reflections in Large Reasoning Models
von: Liu, Kaiyuan, et al.
Veröffentlicht: (2025)
von: Liu, Kaiyuan, et al.
Veröffentlicht: (2025)
CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process
von: Bi, Jinhe, et al.
Veröffentlicht: (2025)
von: Bi, Jinhe, et al.
Veröffentlicht: (2025)
How Likely Do LLMs with CoT Mimic Human Reasoning?
von: Bao, Guangsheng, et al.
Veröffentlicht: (2024)
von: Bao, Guangsheng, et al.
Veröffentlicht: (2024)
Select2Reason: Efficient Instruction-Tuning Data Selection for Long-CoT Reasoning
von: Yang, Cehao, et al.
Veröffentlicht: (2025)
von: Yang, Cehao, et al.
Veröffentlicht: (2025)
RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?
von: Xu, Haotian, et al.
Veröffentlicht: (2025)
von: Xu, Haotian, et al.
Veröffentlicht: (2025)
Generating Effective CoT Traces for Mitigating Causal Hallucination
von: Zhao, Yiheng, et al.
Veröffentlicht: (2026)
von: Zhao, Yiheng, et al.
Veröffentlicht: (2026)
Improve Student's Reasoning Generalizability through Cascading Decomposed CoTs Distillation
von: Dai, Chengwei, et al.
Veröffentlicht: (2024)
von: Dai, Chengwei, et al.
Veröffentlicht: (2024)
Improving Complex Reasoning with Dynamic Prompt Corruption: A soft prompt Optimization Approach
von: Fan, Sinan, et al.
Veröffentlicht: (2025)
von: Fan, Sinan, et al.
Veröffentlicht: (2025)
AS-ES Learning: Towards Efficient CoT Learning in Small Models
von: Xi, Nuwa, et al.
Veröffentlicht: (2024)
von: Xi, Nuwa, et al.
Veröffentlicht: (2024)
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs
von: Chen, Jierun, et al.
Veröffentlicht: (2025)
von: Chen, Jierun, et al.
Veröffentlicht: (2025)
Long or short CoT? Investigating Instance-level Switch of Large Reasoning Models
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2025)
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2025)
Focus on Your Question! Interpreting and Mitigating Toxic CoT Problems in Commonsense Reasoning
von: Li, Jiachun, et al.
Veröffentlicht: (2024)
von: Li, Jiachun, et al.
Veröffentlicht: (2024)
CoT2Align: Cross-Chain of Thought Distillation via Optimal Transport Alignment for Language Models with Different Tokenizers
von: Le, Anh Duc, et al.
Veröffentlicht: (2025)
von: Le, Anh Duc, et al.
Veröffentlicht: (2025)
Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
von: Li, Ang, et al.
Veröffentlicht: (2025)
von: Li, Ang, et al.
Veröffentlicht: (2025)
MIND: From Passive Mimicry to Active Reasoning through Capability-Aware Multi-Perspective CoT Distillation
von: Cui, Jin, et al.
Veröffentlicht: (2026)
von: Cui, Jin, et al.
Veröffentlicht: (2026)
Towards Efficient CoT Distillation: Self-Guided Rationale Selector for Better Performance with Fewer Rationales
von: Yan, Jianzhi, et al.
Veröffentlicht: (2025)
von: Yan, Jianzhi, et al.
Veröffentlicht: (2025)
LatentChem: From Textual CoT to Latent Thinking in Chemical Reasoning
von: Ye, Xinwu, et al.
Veröffentlicht: (2026)
von: Ye, Xinwu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Where Did This Sentence Come From? Tracing Provenance in LLM Reasoning Distillation
von: Liu, Kaiyuan, et al.
Veröffentlicht: (2025) -
On the Step Length Confounding in LLM Reasoning Data Selection
von: Wang, Bing, et al.
Veröffentlicht: (2026) -
Backtracking When It Strays: Mitigating Dual Exposure Biases in LLM Reasoning Distillation
von: Wang, Bing, et al.
Veröffentlicht: (2026) -
Enhancing Chain-of-Thought Reasoning with Critical Representation Fine-tuning
von: Huang, Chenxi, et al.
Veröffentlicht: (2025) -
Don't Take Things Out of Context: Attention Intervention for Enhancing Chain-of-Thought Reasoning in Large Language Models
von: Yan, Shaotian, et al.
Veröffentlicht: (2025)