PAINT: Partial-Solution Adaptive Interpolated Training for Self-Distilled Reasoners
Fuente:
arXiv
Saved in:
| Main Authors: | Tan, Zhiquan, Hong, Yinrong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Supervised On-Policy Distillation for Reasoning Language Models
by: Tan, Zhiquan, et al.
Published: (2026)
by: Tan, Zhiquan, et al.
Published: (2026)
A Theoretical Lens for RL-Tuned Language Models via Energy-Based Models
by: Tan, Zhiquan, et al.
Published: (2025)
by: Tan, Zhiquan, et al.
Published: (2025)
Inference-Cost-Aware Dynamic Tree Construction for Efficient Inference in Large Language Models
by: Hong, Yinrong, et al.
Published: (2025)
by: Hong, Yinrong, et al.
Published: (2025)
Learning from Partial Chain-of-Thought via Truncated-Reasoning Self-Distillation
by: Silvestri, Gianluigi, et al.
Published: (2026)
by: Silvestri, Gianluigi, et al.
Published: (2026)
Information-Theoretic Perspectives on Optimizers
by: Tan, Zhiquan, et al.
Published: (2025)
by: Tan, Zhiquan, et al.
Published: (2025)
Understanding Grokking Through A Robustness Viewpoint
by: Tan, Zhiquan, et al.
Published: (2023)
by: Tan, Zhiquan, et al.
Published: (2023)
Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
by: Zhao, Siyan, et al.
Published: (2026)
by: Zhao, Siyan, et al.
Published: (2026)
Exploring Information-Theoretic Metrics Associated with Neural Collapse in Supervised Training
by: Song, Kun, et al.
Published: (2024)
by: Song, Kun, et al.
Published: (2024)
TAID: Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models
by: Shing, Makoto, et al.
Published: (2025)
by: Shing, Makoto, et al.
Published: (2025)
Partial Distribution Alignment via Adaptive Optimal Transport
by: Yang, Pei, et al.
Published: (2025)
by: Yang, Pei, et al.
Published: (2025)
ThinkSwitch: Context Distillation with LoRA and Weight Interpolation for Specific-Purpose Reasoning Tasks
by: Saini, Dhruv, et al.
Published: (2026)
by: Saini, Dhruv, et al.
Published: (2026)
Boomerang Distillation Enables Zero-Shot Model Size Interpolation
by: Kangaslahti, Sara, et al.
Published: (2025)
by: Kangaslahti, Sara, et al.
Published: (2025)
Rethinking Self-Distillation: Label Averaging and Enhanced Soft Label Refinement with Partial Labels
by: Jeong, Hyeonsu, et al.
Published: (2024)
by: Jeong, Hyeonsu, et al.
Published: (2024)
Can I understand what I create? Self-Knowledge Evaluation of Large Language Models
by: Tan, Zhiquan, et al.
Published: (2024)
by: Tan, Zhiquan, et al.
Published: (2024)
CRISP: Compressed Reasoning via Iterative Self-Policy Distillation
by: Sang, Hejian, et al.
Published: (2026)
by: Sang, Hejian, et al.
Published: (2026)
TED: Training-Free Experience Distillation for Multimodal Reasoning
by: Yuan, Shuozhi, et al.
Published: (2026)
by: Yuan, Shuozhi, et al.
Published: (2026)
Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-Training of Deep Networks
by: Joshi, Siddharth, et al.
Published: (2024)
by: Joshi, Siddharth, et al.
Published: (2024)
How to Train the Teacher Model for Effective Knowledge Distillation
by: Hamidi, Shayan Mohajer, et al.
Published: (2024)
by: Hamidi, Shayan Mohajer, et al.
Published: (2024)
Adaptive Self-Distillation for Minimizing Client Drift in Heterogeneous Federated Learning
by: Yashwanth, M, et al.
Published: (2023)
by: Yashwanth, M, et al.
Published: (2023)
Can Large Reasoning Models Self-Train?
by: Shafayat, Sheikh, et al.
Published: (2025)
by: Shafayat, Sheikh, et al.
Published: (2025)
Annealing Self-Distillation Rectification Improves Adversarial Training
by: Wu, Yu-Yu, et al.
Published: (2023)
by: Wu, Yu-Yu, et al.
Published: (2023)
ATLAS: Adapter-Based Multi-Modal Continual Learning with a Two-Stage Learning Strategy
by: Li, Hong, et al.
Published: (2024)
by: Li, Hong, et al.
Published: (2024)
Matrix Information Theory for Self-Supervised Learning
by: Zhang, Yifan, et al.
Published: (2023)
by: Zhang, Yifan, et al.
Published: (2023)
AdaSTaR: Adaptive Data Sampling for Training Self-Taught Reasoners
by: Koh, Woosung, et al.
Published: (2025)
by: Koh, Woosung, et al.
Published: (2025)
Federated Hybrid Training and Self-Adversarial Distillation: Towards Robust Edge Networks
by: Qiao, Yu, et al.
Published: (2024)
by: Qiao, Yu, et al.
Published: (2024)
Training-Free Generative Modeling via Kernelized Stochastic Interpolants
by: Coeurdoux, Florentin, et al.
Published: (2026)
by: Coeurdoux, Florentin, et al.
Published: (2026)
Interpolated-MLPs: Controllable Inductive Bias
by: Wu, Sean, et al.
Published: (2024)
by: Wu, Sean, et al.
Published: (2024)
Learning While Staying Curious: Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
Data-Driven Self-Supervised Learning for the Discovery of Solution Singularity for Partial Differential Equations
by: Cai, Difeng, et al.
Published: (2025)
by: Cai, Difeng, et al.
Published: (2025)
Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
by: Kim, Jeonghye, et al.
Published: (2026)
by: Kim, Jeonghye, et al.
Published: (2026)
Mitigating Adaptive Attacks against Reasoning Models with Activation Consistency Training
by: Shah, Avidan, et al.
Published: (2026)
by: Shah, Avidan, et al.
Published: (2026)
Adaptive Interpolation-Synthesis for Motion In-Betweening on Keyframe-Based Animation
by: Raël, Anton, et al.
Published: (2026)
by: Raël, Anton, et al.
Published: (2026)
Stable Nonconvex-Nonconcave Training via Linear Interpolation
by: Pethick, Thomas, et al.
Published: (2023)
by: Pethick, Thomas, et al.
Published: (2023)
The Information of Large Language Model Geometry
by: Tan, Zhiquan, et al.
Published: (2024)
by: Tan, Zhiquan, et al.
Published: (2024)
SpecPV: Improving Self-Speculative Decoding for Long-Context Generation via Partial Verification
by: Tan, Zhendong, et al.
Published: (2025)
by: Tan, Zhendong, et al.
Published: (2025)
Osmosis Distillation: Model Hijacking with the Fewest Samples
by: Shi, Yuchen, et al.
Published: (2026)
by: Shi, Yuchen, et al.
Published: (2026)
Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
by: Kim, Jeonghye, et al.
Published: (2026)
by: Kim, Jeonghye, et al.
Published: (2026)
When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning
by: Liu, Xiaogeng, et al.
Published: (2026)
by: Liu, Xiaogeng, et al.
Published: (2026)
ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains
by: Zhao, Ziqi, et al.
Published: (2026)
by: Zhao, Ziqi, et al.
Published: (2026)
Restoring the Sweet Spot: Pass-Rate Weighted Self-Distillation for LLM Reasoning
by: Liu, Zehao, et al.
Published: (2026)
by: Liu, Zehao, et al.
Published: (2026)
Similar Items
-
Self-Supervised On-Policy Distillation for Reasoning Language Models
by: Tan, Zhiquan, et al.
Published: (2026) -
A Theoretical Lens for RL-Tuned Language Models via Energy-Based Models
by: Tan, Zhiquan, et al.
Published: (2025) -
Inference-Cost-Aware Dynamic Tree Construction for Efficient Inference in Large Language Models
by: Hong, Yinrong, et al.
Published: (2025) -
Learning from Partial Chain-of-Thought via Truncated-Reasoning Self-Distillation
by: Silvestri, Gianluigi, et al.
Published: (2026) -
Information-Theoretic Perspectives on Optimizers
by: Tan, Zhiquan, et al.
Published: (2025)