The Illusion of Certainty: Decoupling Capability and Calibration in On-Policy Distillation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Jiaxin, Peng, Xiangyu, Chen, Qinglin, Ye, Qinyuan, Xiong, Caiming, Wu, Chien-Sheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Agentic Confidence Calibration
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2026)
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2026)
Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
von: Xu, Austin, et al.
Veröffentlicht: (2025)
von: Xu, Austin, et al.
Veröffentlicht: (2025)
Capturing LLM Capabilities via Evidence-Calibrated Query Clustering
von: Wu, Fangzhou, et al.
Veröffentlicht: (2026)
von: Wu, Fangzhou, et al.
Veröffentlicht: (2026)
Calibration-Aware Policy Optimization for Reasoning LLMs
von: Wang, Ziqi, et al.
Veröffentlicht: (2026)
von: Wang, Ziqi, et al.
Veröffentlicht: (2026)
GRAFT: Decoupling Ranking and Calibration for Survival Analysis
von: Ashhad, Mohammad, et al.
Veröffentlicht: (2026)
von: Ashhad, Mohammad, et al.
Veröffentlicht: (2026)
$\boldsymbol{f}$-OPD: Stabilizing Long-Horizon On-Policy Distillation with Freshness-Aware Control
von: Chen, Xianwei, et al.
Veröffentlicht: (2026)
von: Chen, Xianwei, et al.
Veröffentlicht: (2026)
Validity-Calibrated Reasoning Distillation
von: Saadi, Khouloud, et al.
Veröffentlicht: (2026)
von: Saadi, Khouloud, et al.
Veröffentlicht: (2026)
Extreme Region Policy Distillation
von: Chen, Changyu, et al.
Veröffentlicht: (2026)
von: Chen, Changyu, et al.
Veröffentlicht: (2026)
Dirichlet-Based Prediction Calibration for Learning with Noisy Labels
von: Zong, Chen-Chen, et al.
Veröffentlicht: (2024)
von: Zong, Chen-Chen, et al.
Veröffentlicht: (2024)
Distillation Traps and Guards: A Calibration Knob for LLM Distillability
von: Zhan, Weixiao, et al.
Veröffentlicht: (2026)
von: Zhan, Weixiao, et al.
Veröffentlicht: (2026)
Pretrain Value, Not Reward: Decoupled Value Policy Optimization
von: Huang, Chenghua, et al.
Veröffentlicht: (2025)
von: Huang, Chenghua, et al.
Veröffentlicht: (2025)
On Calibration of Large Language Models: From Response To Capability
von: Yang, Sin-Han, et al.
Veröffentlicht: (2026)
von: Yang, Sin-Han, et al.
Veröffentlicht: (2026)
Enabling High Data Throughput Reinforcement Learning on GPUs: A Domain Agnostic Framework for Data-Driven Scientific Research
von: Lan, Tian, et al.
Veröffentlicht: (2024)
von: Lan, Tian, et al.
Veröffentlicht: (2024)
DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization
von: Li, Gang, et al.
Veröffentlicht: (2025)
von: Li, Gang, et al.
Veröffentlicht: (2025)
Are Flat Minima an Illusion?
von: Bennett, Michael Timothy
Veröffentlicht: (2026)
von: Bennett, Michael Timothy
Veröffentlicht: (2026)
SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting
von: Zheng, Binbin, et al.
Veröffentlicht: (2026)
von: Zheng, Binbin, et al.
Veröffentlicht: (2026)
Improving Prediction Certainty Estimation for Reliable Early Exiting via Null Space Projection
von: He, Jianing, et al.
Veröffentlicht: (2025)
von: He, Jianing, et al.
Veröffentlicht: (2025)
The Over-Certainty Phenomenon in Modern Test-Time Adaptation Algorithms
von: Amin, Fin, et al.
Veröffentlicht: (2024)
von: Amin, Fin, et al.
Veröffentlicht: (2024)
Proximal Policy Distillation
von: Spigler, Giacomo
Veröffentlicht: (2024)
von: Spigler, Giacomo
Veröffentlicht: (2024)
Unleashing the Denoising Capability of Diffusion Prior for Solving Inverse Problems
von: Zhang, Jiawei, et al.
Veröffentlicht: (2024)
von: Zhang, Jiawei, et al.
Veröffentlicht: (2024)
Function Induction and Task Generalization: An Interpretability Study with Off-by-One Addition
von: Ye, Qinyuan, et al.
Veröffentlicht: (2025)
von: Ye, Qinyuan, et al.
Veröffentlicht: (2025)
Stress-Testing Long-Context Language Models with Lifelong ICL and Task Haystack
von: Xu, Xiaoyue, et al.
Veröffentlicht: (2024)
von: Xu, Xiaoyue, et al.
Veröffentlicht: (2024)
Towards Flash Thinking via Decoupled Advantage Policy Optimization
von: Tan, Zezhong, et al.
Veröffentlicht: (2025)
von: Tan, Zezhong, et al.
Veröffentlicht: (2025)
When Maximum Entropy Misleads Policy Optimization
von: Zhang, Ruipeng, et al.
Veröffentlicht: (2025)
von: Zhang, Ruipeng, et al.
Veröffentlicht: (2025)
OPD+: Rethinking the Advantage Design for On-Policy Distillation
von: Zhao, Hanyang, et al.
Veröffentlicht: (2026)
von: Zhao, Hanyang, et al.
Veröffentlicht: (2026)
ADWIN: Adaptive Windows for Horizon-Aware On-Policy Distillation
von: Liang, Kun, et al.
Veröffentlicht: (2026)
von: Liang, Kun, et al.
Veröffentlicht: (2026)
AlignDistil: Token-Level Language Model Alignment as Adaptive Policy Distillation
von: Zhang, Songming, et al.
Veröffentlicht: (2025)
von: Zhang, Songming, et al.
Veröffentlicht: (2025)
PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2025)
Dynamic Evidence Decoupling for Trusted Multi-view Learning
von: Liu, Ying, et al.
Veröffentlicht: (2024)
von: Liu, Ying, et al.
Veröffentlicht: (2024)
HDPO: Hybrid Distillation Policy Optimization via Privileged Self-Distillation
von: Ding, Ken
Veröffentlicht: (2026)
von: Ding, Ken
Veröffentlicht: (2026)
TIP: Token Importance in On-Policy Distillation
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
Online Policy Distillation with Decision-Attention
von: Yu, Xinqiang, et al.
Veröffentlicht: (2024)
von: Yu, Xinqiang, et al.
Veröffentlicht: (2024)
Curriculum Learning for Efficient Chain-of-Thought Distillation via Structure-Aware Masking and GRPO
von: Yu, Bowen, et al.
Veröffentlicht: (2026)
von: Yu, Bowen, et al.
Veröffentlicht: (2026)
An approach of deep reinforcement learning for maximizing the net present value of stochastic projects
von: Xu, Wei, et al.
Veröffentlicht: (2025)
von: Xu, Wei, et al.
Veröffentlicht: (2025)
Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level
von: Jia, Nan, et al.
Veröffentlicht: (2026)
von: Jia, Nan, et al.
Veröffentlicht: (2026)
Context Distillation as Latent Memory Management
von: Zheng, Ziyang, et al.
Veröffentlicht: (2026)
von: Zheng, Ziyang, et al.
Veröffentlicht: (2026)
Constraint Decoupled Latent Diffusion for Protein Backmapping
von: Han, Xu, et al.
Veröffentlicht: (2024)
von: Han, Xu, et al.
Veröffentlicht: (2024)
Evidentially Calibrated Source-Free Time-Series Domain Adaptation with Temporal Imputation
von: Ragab, Mohamed, et al.
Veröffentlicht: (2024)
von: Ragab, Mohamed, et al.
Veröffentlicht: (2024)
The Illusion of Readiness in Health AI
von: Gu, Yu, et al.
Veröffentlicht: (2025)
von: Gu, Yu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Agentic Confidence Calibration
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2026) -
Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
von: Xu, Austin, et al.
Veröffentlicht: (2025) -
Capturing LLM Capabilities via Evidence-Calibrated Query Clustering
von: Wu, Fangzhou, et al.
Veröffentlicht: (2026) -
Calibration-Aware Policy Optimization for Reasoning LLMs
von: Wang, Ziqi, et al.
Veröffentlicht: (2026) -
GRAFT: Decoupling Ranking and Calibration for Survival Analysis
von: Ashhad, Mohammad, et al.
Veröffentlicht: (2026)