Trust Region On-Policy Distillation
Fuente:
arXiv
Guardado en:
| Autores principales: | Xing, Xingrun, Wang, Haoqing, Gao, Boyan, Li, Ziheng, Tang, Yehui |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
When Do LLMs Reason? A Dynamical Systems View via Entropy Phase Transitions
por: Xia, Wei, et al.
Publicado: (2026)
por: Xia, Wei, et al.
Publicado: (2026)
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
por: Long, Xiang, et al.
Publicado: (2026)
por: Long, Xiang, et al.
Publicado: (2026)
SpikeLLM: Scaling up Spiking Neural Network to Large Language Models via Saliency-based Spiking
por: Xing, Xingrun, et al.
Publicado: (2024)
por: Xing, Xingrun, et al.
Publicado: (2024)
TRE: Encouraging Exploration in the Trust Region
por: Huang, Chao, et al.
Publicado: (2026)
por: Huang, Chao, et al.
Publicado: (2026)
MLKD-BERT: Multi-level Knowledge Distillation for Pre-trained Language Models
por: Zhang, Ying, et al.
Publicado: (2024)
por: Zhang, Ying, et al.
Publicado: (2024)
PocketLLM: Ultimate Compression of Large Language Models via Meta Networks
por: Tian, Ye, et al.
Publicado: (2025)
por: Tian, Ye, et al.
Publicado: (2025)
Mixture of Lookup Experts
por: Jie, Shibo, et al.
Publicado: (2025)
por: Jie, Shibo, et al.
Publicado: (2025)
The Extrapolation Cliff in On-Policy Distillation of Near-Deterministic Structured Outputs
por: Li, Xin, et al.
Publicado: (2026)
por: Li, Xin, et al.
Publicado: (2026)
Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning
por: Yang, Xuewei, et al.
Publicado: (2026)
por: Yang, Xuewei, et al.
Publicado: (2026)
Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing
por: Rang, Miao, et al.
Publicado: (2026)
por: Rang, Miao, et al.
Publicado: (2026)
Kangaroo: Lossless Self-Speculative Decoding via Double Early Exiting
por: Liu, Fangcheng, et al.
Publicado: (2024)
por: Liu, Fangcheng, et al.
Publicado: (2024)
PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization
por: Rahman, Ben
Publicado: (2025)
por: Rahman, Ben
Publicado: (2025)
Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
por: Zhao, Siyan, et al.
Publicado: (2026)
por: Zhao, Siyan, et al.
Publicado: (2026)
Sinkhorn Distance Minimization for Knowledge Distillation
por: Cui, Xiao, et al.
Publicado: (2024)
por: Cui, Xiao, et al.
Publicado: (2024)
TRAM: Bridging Trust Regions and Sharpness Aware Minimization
por: Sherborne, Tom, et al.
Publicado: (2023)
por: Sherborne, Tom, et al.
Publicado: (2023)
DenseMamba: State Space Models with Dense Hidden Connection for Efficient Large Language Models
por: He, Wei, et al.
Publicado: (2024)
por: He, Wei, et al.
Publicado: (2024)
SpikeLM: Towards General Spike-Driven Language Modeling via Elastic Bi-Spiking Mechanisms
por: Xing, Xingrun, et al.
Publicado: (2024)
por: Xing, Xingrun, et al.
Publicado: (2024)
SODA: Semi On-Policy Black-Box Distillation for Large Language Models
por: Chen, Xiwen, et al.
Publicado: (2026)
por: Chen, Xiwen, et al.
Publicado: (2026)
Mitigating Training Imbalance in LLM Fine-Tuning via Selective Parameter Merging
por: Ju, Yiming, et al.
Publicado: (2024)
por: Ju, Yiming, et al.
Publicado: (2024)
A Survey of On-Policy Distillation for Large Language Models
por: Song, Mingyang, et al.
Publicado: (2026)
por: Song, Mingyang, et al.
Publicado: (2026)
Scaling Reasoning Efficiently via Relaxed On-Policy Distillation
por: Ko, Jongwoo, et al.
Publicado: (2026)
por: Ko, Jongwoo, et al.
Publicado: (2026)
OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification
por: Zhou, Yuhang, et al.
Publicado: (2026)
por: Zhou, Yuhang, et al.
Publicado: (2026)
Efficient Second-Order Neural Network Optimization via Adaptive Trust Region Methods
por: Vo, James
Publicado: (2024)
por: Vo, James
Publicado: (2024)
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
por: Li, Yaxuan, et al.
Publicado: (2026)
por: Li, Yaxuan, et al.
Publicado: (2026)
SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting
por: Zheng, Binbin, et al.
Publicado: (2026)
por: Zheng, Binbin, et al.
Publicado: (2026)
Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient Reasoners
por: Xu, Xin, et al.
Publicado: (2025)
por: Xu, Xin, et al.
Publicado: (2025)
ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains
por: Zhao, Ziqi, et al.
Publicado: (2026)
por: Zhao, Ziqi, et al.
Publicado: (2026)
AlignDistil: Token-Level Language Model Alignment as Adaptive Policy Distillation
por: Zhang, Songming, et al.
Publicado: (2025)
por: Zhang, Songming, et al.
Publicado: (2025)
ORPO-Distill: Mixed-Policy Preference Optimization for Cross-Architecture LLM Distillation
por: Singh, Aasheesh, et al.
Publicado: (2025)
por: Singh, Aasheesh, et al.
Publicado: (2025)
BitNet Distillation
por: Wu, Xun, et al.
Publicado: (2025)
por: Wu, Xun, et al.
Publicado: (2025)
CBQ: Cross-Block Quantization for Large Language Models
por: Ding, Xin, et al.
Publicado: (2023)
por: Ding, Xin, et al.
Publicado: (2023)
Memory-Space Visual Prompting for Efficient Vision-Language Fine-Tuning
por: Jie, Shibo, et al.
Publicado: (2024)
por: Jie, Shibo, et al.
Publicado: (2024)
Scaling over Scaling: Exploring Test-Time Scaling Plateau in Large Reasoning Models
por: Wang, Jian, et al.
Publicado: (2025)
por: Wang, Jian, et al.
Publicado: (2025)
MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate
por: Wang, Jianze, et al.
Publicado: (2026)
por: Wang, Jianze, et al.
Publicado: (2026)
Self-Distilled RLVR
por: Yang, Chenxu, et al.
Publicado: (2026)
por: Yang, Chenxu, et al.
Publicado: (2026)
A Survey on Transformer Compression
por: Tang, Yehui, et al.
Publicado: (2024)
por: Tang, Yehui, et al.
Publicado: (2024)
BWArea Model: Learning World Model, Inverse Dynamics, and Policy for Controllable Language Generation
por: Jia, Chengxing, et al.
Publicado: (2024)
por: Jia, Chengxing, et al.
Publicado: (2024)
Outcome-Grounded Advantage Reshaping for Fine-Grained Credit Assignment in Mathematical Reasoning
por: Li, Ziheng, et al.
Publicado: (2026)
por: Li, Ziheng, et al.
Publicado: (2026)
Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
por: Fu, Yuqian, et al.
Publicado: (2026)
por: Fu, Yuqian, et al.
Publicado: (2026)
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
por: Yang, Wenkai, et al.
Publicado: (2026)
por: Yang, Wenkai, et al.
Publicado: (2026)
Ejemplares similares
-
When Do LLMs Reason? A Dynamical Systems View via Entropy Phase Transitions
por: Xia, Wei, et al.
Publicado: (2026) -
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
por: Long, Xiang, et al.
Publicado: (2026) -
SpikeLLM: Scaling up Spiking Neural Network to Large Language Models via Saliency-based Spiking
por: Xing, Xingrun, et al.
Publicado: (2024) -
TRE: Encouraging Exploration in the Trust Region
por: Huang, Chao, et al.
Publicado: (2026) -
MLKD-BERT: Multi-level Knowledge Distillation for Pre-trained Language Models
por: Zhang, Ying, et al.
Publicado: (2024)