Validity-Calibrated Reasoning Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Saadi, Khouloud, Wang, Di |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dissecting Representation Misalignment in Contrastive Learning via Influence Function
by: Hu, Lijie, et al.
Published: (2024)
by: Hu, Lijie, et al.
Published: (2024)
Distillation Traps and Guards: A Calibration Knob for LLM Distillability
by: Zhan, Weixiao, et al.
Published: (2026)
by: Zhan, Weixiao, et al.
Published: (2026)
The Illusion of Certainty: Decoupling Capability and Calibration in On-Policy Distillation
by: Zhang, Jiaxin, et al.
Published: (2026)
by: Zhang, Jiaxin, et al.
Published: (2026)
Calibration-Aware Policy Optimization for Reasoning LLMs
by: Wang, Ziqi, et al.
Published: (2026)
by: Wang, Ziqi, et al.
Published: (2026)
TED: Training-Free Experience Distillation for Multimodal Reasoning
by: Yuan, Shuozhi, et al.
Published: (2026)
by: Yuan, Shuozhi, et al.
Published: (2026)
Self-Attentive Spatio-Temporal Calibration for Precise Intermediate Layer Matching in ANN-to-SNN Distillation
by: Hong, Di, et al.
Published: (2025)
by: Hong, Di, et al.
Published: (2025)
HEAL: Hindsight Entropy-Assisted Learning for Reasoning Distillation
by: Zhang, Wenjing, et al.
Published: (2026)
by: Zhang, Wenjing, et al.
Published: (2026)
VISTA: Validation-Informed Trajectory Adaptation via Self-Distillation
by: Corn, Eli, et al.
Published: (2026)
by: Corn, Eli, et al.
Published: (2026)
Reinforcement-aware Knowledge Distillation for LLM Reasoning
by: Zhang, Zhaoyang, et al.
Published: (2026)
by: Zhang, Zhaoyang, et al.
Published: (2026)
Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning
by: Yang, Zhicheng, et al.
Published: (2026)
by: Yang, Zhicheng, et al.
Published: (2026)
Distilling Calibration via Conformalized Credal Inference
by: Huang, Jiayi, et al.
Published: (2025)
by: Huang, Jiayi, et al.
Published: (2025)
The Quest for Efficient Reasoning: A Data-Centric Benchmark to CoT Distillation
by: Zhang, Ruichen, et al.
Published: (2025)
by: Zhang, Ruichen, et al.
Published: (2025)
Fast and Effective On-policy Distillation from Reasoning Prefixes
by: Zhang, Dongxu, et al.
Published: (2026)
by: Zhang, Dongxu, et al.
Published: (2026)
When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning
by: Liu, Xiaogeng, et al.
Published: (2026)
by: Liu, Xiaogeng, et al.
Published: (2026)
Enhanced geometry prediction in laser directed energy deposition using meta-learning
by: Saadi, Abdul Malik Al Mardhouf Al, et al.
Published: (2025)
by: Saadi, Abdul Malik Al Mardhouf Al, et al.
Published: (2025)
OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning
by: Yang, Yuxiao, et al.
Published: (2026)
by: Yang, Yuxiao, et al.
Published: (2026)
Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning
by: Wu, Xiaojun, et al.
Published: (2025)
by: Wu, Xiaojun, et al.
Published: (2025)
Hindsight Hint Distillation: Scaffolded Reasoning for SWE Agents from CoT-free Answers
by: Wang, Shengjie, et al.
Published: (2026)
by: Wang, Shengjie, et al.
Published: (2026)
Learning to Reason: Temporal Saliency Distillation for Interpretable Knowledge Transfer
by: Dehigahawattage, Nilushika Udayangani Hewa, et al.
Published: (2026)
by: Dehigahawattage, Nilushika Udayangani Hewa, et al.
Published: (2026)
The Role of Teacher Calibration in Knowledge Distillation
by: Kim, Suyoung, et al.
Published: (2025)
by: Kim, Suyoung, et al.
Published: (2025)
From Entropy to Calibrated Uncertainty: Training Language Models to Reason About Uncertainty
by: Jenane, Azza, et al.
Published: (2026)
by: Jenane, Azza, et al.
Published: (2026)
Distillation Contrastive Decoding: Improving LLMs Reasoning with Contrastive Decoding and Distillation
by: Phan, Phuc, et al.
Published: (2024)
by: Phan, Phuc, et al.
Published: (2024)
Learning from Partial Chain-of-Thought via Truncated-Reasoning Self-Distillation
by: Silvestri, Gianluigi, et al.
Published: (2026)
by: Silvestri, Gianluigi, et al.
Published: (2026)
Restoring the Sweet Spot: Pass-Rate Weighted Self-Distillation for LLM Reasoning
by: Liu, Zehao, et al.
Published: (2026)
by: Liu, Zehao, et al.
Published: (2026)
Structural Rationale Distillation via Reasoning Space Compression
by: Yang, Jialin, et al.
Published: (2026)
by: Yang, Jialin, et al.
Published: (2026)
Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation
by: Kim, Juno, et al.
Published: (2025)
by: Kim, Juno, et al.
Published: (2025)
Where Paths Split: Localized, Calibrated Control of Moral Reasoning in Large Language Models
by: Yuan, Chenchen, et al.
Published: (2026)
by: Yuan, Chenchen, et al.
Published: (2026)
CarBoN: Calibrated Best-of-N Sampling Improves Test-time Reasoning
by: Tang, Yung-Chen, et al.
Published: (2025)
by: Tang, Yung-Chen, et al.
Published: (2025)
Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation
by: Wu, Yecheng, et al.
Published: (2026)
by: Wu, Yecheng, et al.
Published: (2026)
SpatialTraceGen: High-Fidelity Traces for Efficient VLM Spatial Reasoning Distillation
by: Huh, Gio, et al.
Published: (2025)
by: Huh, Gio, et al.
Published: (2025)
SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting
by: Zheng, Binbin, et al.
Published: (2026)
by: Zheng, Binbin, et al.
Published: (2026)
What Should Feature Distillation Transfer in LLMs? A Task-Tangent Geometry View
by: Saadi, Khouloud, et al.
Published: (2025)
by: Saadi, Khouloud, et al.
Published: (2025)
PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence
by: Xu, Yuanda, et al.
Published: (2026)
by: Xu, Yuanda, et al.
Published: (2026)
Geometric Analysis of Reasoning Trajectories: A Phase Space Approach to Understanding Valid and Invalid Multi-Hop Reasoning in LLMs
by: Marin, Javier
Published: (2024)
by: Marin, Javier
Published: (2024)
SAGE-32B: Agentic Reasoning via Iterative Distillation
by: Jha, Basab, et al.
Published: (2026)
by: Jha, Basab, et al.
Published: (2026)
ThinkSwitch: Context Distillation with LoRA and Weight Interpolation for Specific-Purpose Reasoning Tasks
by: Saini, Dhruv, et al.
Published: (2026)
by: Saini, Dhruv, et al.
Published: (2026)
Online Reasoning Calibration: Test-Time Training Enables Generalizable Conformal LLM Reasoning
by: Zhou, Cai, et al.
Published: (2026)
by: Zhou, Cai, et al.
Published: (2026)
Reinforced Reasoning for Embodied Planning
by: Wu, Di, et al.
Published: (2025)
by: Wu, Di, et al.
Published: (2025)
Distilling Autoregressive Models to Obtain High-Performance Non-Autoregressive Solvers for Vehicle Routing Problems with Faster Inference Speed
by: Xiao, Yubin, et al.
Published: (2023)
by: Xiao, Yubin, et al.
Published: (2023)
Extreme Region Policy Distillation
by: Chen, Changyu, et al.
Published: (2026)
by: Chen, Changyu, et al.
Published: (2026)
Similar Items
-
Dissecting Representation Misalignment in Contrastive Learning via Influence Function
by: Hu, Lijie, et al.
Published: (2024) -
Distillation Traps and Guards: A Calibration Knob for LLM Distillability
by: Zhan, Weixiao, et al.
Published: (2026) -
The Illusion of Certainty: Decoupling Capability and Calibration in On-Policy Distillation
by: Zhang, Jiaxin, et al.
Published: (2026) -
Calibration-Aware Policy Optimization for Reasoning LLMs
by: Wang, Ziqi, et al.
Published: (2026) -
TED: Training-Free Experience Distillation for Multimodal Reasoning
by: Yuan, Shuozhi, et al.
Published: (2026)