CRISP: Compressed Reasoning via Iterative Self-Policy Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Sang, Hejian, Xu, Yuanda, Zhou, Zhengze, He, Ran, Wang, Zhipeng, Sun, Jiachen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence
by: Xu, Yuanda, et al.
Published: (2026)
by: Xu, Yuanda, et al.
Published: (2026)
TIP: Token Importance in On-Policy Distillation
by: Xu, Yuanda, et al.
Published: (2026)
by: Xu, Yuanda, et al.
Published: (2026)
Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training
by: Xu, Yuanda, et al.
Published: (2026)
by: Xu, Yuanda, et al.
Published: (2026)
Overconfident Errors Need Stronger Correction: Asymmetric Confidence Penalties for Reinforcement Learning
by: Xu, Yuanda, et al.
Published: (2026)
by: Xu, Yuanda, et al.
Published: (2026)
Not all tokens are needed(NAT): token efficient reinforcement learning
by: Sang, Hejian, et al.
Published: (2026)
by: Sang, Hejian, et al.
Published: (2026)
Robust Batch-Level Query Routing for Large Language Models under Cost and Capacity Constraints
by: Markovic-Voronov, Jelena, et al.
Published: (2026)
by: Markovic-Voronov, Jelena, et al.
Published: (2026)
SODA: Semi On-Policy Black-Box Distillation for Large Language Models
by: Chen, Xiwen, et al.
Published: (2026)
by: Chen, Xiwen, et al.
Published: (2026)
Debunk the Myth of SFT Generalization
by: Lin, Xiaofeng, et al.
Published: (2025)
by: Lin, Xiaofeng, et al.
Published: (2025)
Aligning Diffusion Language Models via Unpaired Preference Optimization
by: Jindal, Vaibhav, et al.
Published: (2025)
by: Jindal, Vaibhav, et al.
Published: (2025)
Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning
by: Yang, Xuewei, et al.
Published: (2026)
by: Yang, Xuewei, et al.
Published: (2026)
Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
by: Zhao, Siyan, et al.
Published: (2026)
by: Zhao, Siyan, et al.
Published: (2026)
Structural Rationale Distillation via Reasoning Space Compression
by: Yang, Jialin, et al.
Published: (2026)
by: Yang, Jialin, et al.
Published: (2026)
Self-Supervised On-Policy Distillation for Reasoning Language Models
by: Tan, Zhiquan, et al.
Published: (2026)
by: Tan, Zhiquan, et al.
Published: (2026)
Reducing the Safety Tax in LLM Safety Alignment with On-Policy Self-Distillation
by: Fu, Yu, et al.
Published: (2026)
by: Fu, Yu, et al.
Published: (2026)
Self Distillation via Iterative Constructive Perturbations
by: Dave, Maheak, et al.
Published: (2025)
by: Dave, Maheak, et al.
Published: (2025)
SAGE-32B: Agentic Reasoning via Iterative Distillation
by: Jha, Basab, et al.
Published: (2026)
by: Jha, Basab, et al.
Published: (2026)
ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains
by: Zhao, Ziqi, et al.
Published: (2026)
by: Zhao, Ziqi, et al.
Published: (2026)
KaVa: Latent Reasoning via Compressed KV-Cache Distillation
by: Kuzina, Anna, et al.
Published: (2025)
by: Kuzina, Anna, et al.
Published: (2025)
Scaling In-Context Online Learning Capability of LLMs via Cross-Episode Meta-RL
by: Lin, Xiaofeng, et al.
Published: (2026)
by: Lin, Xiaofeng, et al.
Published: (2026)
When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning
by: Liu, Xiaogeng, et al.
Published: (2026)
by: Liu, Xiaogeng, et al.
Published: (2026)
OISD: On-Policy Internal Self-Distillation of Language Models
by: Liu, Xinyu, et al.
Published: (2026)
by: Liu, Xinyu, et al.
Published: (2026)
HDPO: Hybrid Distillation Policy Optimization via Privileged Self-Distillation
by: Ding, Ken
Published: (2026)
by: Ding, Ken
Published: (2026)
Scaling Reasoning Efficiently via Relaxed On-Policy Distillation
by: Ko, Jongwoo, et al.
Published: (2026)
by: Ko, Jongwoo, et al.
Published: (2026)
OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning
by: Yang, Yuxiao, et al.
Published: (2026)
by: Yang, Yuxiao, et al.
Published: (2026)
Semi-Supervised Reward Modeling via Iterative Self-Training
by: He, Yifei, et al.
Published: (2024)
by: He, Yifei, et al.
Published: (2024)
Iterative Layer-wise Distillation for Efficient Compression of Large Language Models
by: Kovalev, Grigory, et al.
Published: (2025)
by: Kovalev, Grigory, et al.
Published: (2025)
Examining spatial heterogeneity of ridesourcing demand determinants with explainable machine learning
by: Zhang, Xiaojian, et al.
Published: (2022)
by: Zhang, Xiaojian, et al.
Published: (2022)
Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning
by: Yang, Zhicheng, et al.
Published: (2026)
by: Yang, Zhicheng, et al.
Published: (2026)
One-Step Diffusion Policy: Fast Visuomotor Policies via Diffusion Distillation
by: Wang, Zhendong, et al.
Published: (2024)
by: Wang, Zhendong, et al.
Published: (2024)
Prompting Policies for Multi-step Reasoning and Tool-Use in Black-box LLMs with Iterative Distillation of Experience
by: Sayana, Krishna, et al.
Published: (2026)
by: Sayana, Krishna, et al.
Published: (2026)
Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient Reasoners
by: Xu, Xin, et al.
Published: (2025)
by: Xu, Xin, et al.
Published: (2025)
Distilling the Essence: Efficient Reasoning Distillation via Sequence Truncation
by: Chen, Wei-Rui, et al.
Published: (2025)
by: Chen, Wei-Rui, et al.
Published: (2025)
Self-Distillation as a Performance Recovery Mechanism for LLMs: Counteracting Compression and Catastrophic Forgetting
by: Liu, Chi, et al.
Published: (2026)
by: Liu, Chi, et al.
Published: (2026)
TRACE: Distilling Where It Matters via Token-Routed Self On-Policy Alignment
by: Wang, Jiaxuan, et al.
Published: (2026)
by: Wang, Jiaxuan, et al.
Published: (2026)
From Few to Many: Self-Improving Many-Shot Reasoners Through Iterative Optimization and Generation
by: Wan, Xingchen, et al.
Published: (2025)
by: Wan, Xingchen, et al.
Published: (2025)
Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing
by: Li, Gengsheng, et al.
Published: (2026)
by: Li, Gengsheng, et al.
Published: (2026)
Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing
by: Rang, Miao, et al.
Published: (2026)
by: Rang, Miao, et al.
Published: (2026)
Score and Distribution Matching Policy: Advanced Accelerated Visuomotor Policies via Matched Distillation
by: Jia, Bofang, et al.
Published: (2024)
by: Jia, Bofang, et al.
Published: (2024)
CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning
by: Li, Ran, et al.
Published: (2026)
by: Li, Ran, et al.
Published: (2026)
Iterative Refinement of Flow Policies in Probability Space for Online Reinforcement Learning
by: Sun, Mingyang, et al.
Published: (2025)
by: Sun, Mingyang, et al.
Published: (2025)
Similar Items
-
PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence
by: Xu, Yuanda, et al.
Published: (2026) -
TIP: Token Importance in On-Policy Distillation
by: Xu, Yuanda, et al.
Published: (2026) -
Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training
by: Xu, Yuanda, et al.
Published: (2026) -
Overconfident Errors Need Stronger Correction: Asymmetric Confidence Penalties for Reinforcement Learning
by: Xu, Yuanda, et al.
Published: (2026) -
Not all tokens are needed(NAT): token efficient reinforcement learning
by: Sang, Hejian, et al.
Published: (2026)