Stable On-Policy Distillation through Adaptive Target Reformulation
Fuente:
arXiv
Salvato in:
| Autori principali: | Jang, Ijun, Yeom, Jewon, Yeo, Juan, Lim, Hyunggu, Kim, Taesup |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Efficient Epistemic Uncertainty Estimation for Large Language Models via Knowledge Distillation
di: Park, Seonghyeon, et al.
Pubblicazione: (2026)
di: Park, Seonghyeon, et al.
Pubblicazione: (2026)
Beyond Case Law: Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal QA
di: Chae, Kyubyung, et al.
Pubblicazione: (2026)
di: Chae, Kyubyung, et al.
Pubblicazione: (2026)
Robust Domain Generalization under Divergent Marginal and Conditional Distributions
di: Yeom, Jewon, et al.
Pubblicazione: (2026)
di: Yeom, Jewon, et al.
Pubblicazione: (2026)
Patch-Level Kernel Alignment for Dense Self-Supervised Learning
di: Yeo, Juan, et al.
Pubblicazione: (2025)
di: Yeo, Juan, et al.
Pubblicazione: (2025)
Towards Robust Real-World Multivariate Time Series Forecasting: A Unified Framework for Dependency, Asynchrony, and Missingness
di: Jang, Jinkwan, et al.
Pubblicazione: (2025)
di: Jang, Jinkwan, et al.
Pubblicazione: (2025)
X-PEFT: eXtremely Parameter-Efficient Fine-Tuning for Extreme Multi-Profile Scenarios
di: Kwak, Namju, et al.
Pubblicazione: (2024)
di: Kwak, Namju, et al.
Pubblicazione: (2024)
PiCa: Parameter-Efficient Fine-Tuning with Column Space Projection
di: Hwang, Junseo, et al.
Pubblicazione: (2025)
di: Hwang, Junseo, et al.
Pubblicazione: (2025)
Contrastive Residual Energy Test-time Adaptation
di: Han, Yewon, et al.
Pubblicazione: (2025)
di: Han, Yewon, et al.
Pubblicazione: (2025)
Reset & Distill: A Recipe for Overcoming Negative Transfer in Continual Reinforcement Learning
di: Ahn, Hongjoon, et al.
Pubblicazione: (2024)
di: Ahn, Hongjoon, et al.
Pubblicazione: (2024)
Learning to Act Robustly with View-Invariant Latent Actions
di: Jeong, Youngjoon, et al.
Pubblicazione: (2026)
di: Jeong, Youngjoon, et al.
Pubblicazione: (2026)
ADWIN: Adaptive Windows for Horizon-Aware On-Policy Distillation
di: Liang, Kun, et al.
Pubblicazione: (2026)
di: Liang, Kun, et al.
Pubblicazione: (2026)
HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agents
di: Yeo, Woongyeng, et al.
Pubblicazione: (2026)
di: Yeo, Woongyeng, et al.
Pubblicazione: (2026)
One-Step Generative Policies with Q-Learning: A Reformulation of MeanFlow
di: Wang, Zeyuan, et al.
Pubblicazione: (2025)
di: Wang, Zeyuan, et al.
Pubblicazione: (2025)
PCoreSet: Effective Active Learning through Knowledge Distillation from Vision-Language Models
di: Kang, Seongjae, et al.
Pubblicazione: (2025)
di: Kang, Seongjae, et al.
Pubblicazione: (2025)
An Adversarial Learning Approach to Irregular Time-Series Forecasting
di: Nam, Heejeong, et al.
Pubblicazione: (2024)
di: Nam, Heejeong, et al.
Pubblicazione: (2024)
PRISP: Privacy-Safe Few-Shot Personalization via Lightweight Adaptation
di: Park, Junho, et al.
Pubblicazione: (2026)
di: Park, Junho, et al.
Pubblicazione: (2026)
AlignDistil: Token-Level Language Model Alignment as Adaptive Policy Distillation
di: Zhang, Songming, et al.
Pubblicazione: (2025)
di: Zhang, Songming, et al.
Pubblicazione: (2025)
Proximal Policy Distillation
di: Spigler, Giacomo
Pubblicazione: (2024)
di: Spigler, Giacomo
Pubblicazione: (2024)
On the Internal Representations of Graph Metanetworks
di: Yeom, Taesun, et al.
Pubblicazione: (2025)
di: Yeom, Taesun, et al.
Pubblicazione: (2025)
Extreme Region Policy Distillation
di: Chen, Changyu, et al.
Pubblicazione: (2026)
di: Chen, Changyu, et al.
Pubblicazione: (2026)
Attention-space Contrastive Guidance for Efficient Hallucination Mitigation in LVLMs
di: Jo, Yujin, et al.
Pubblicazione: (2026)
di: Jo, Yujin, et al.
Pubblicazione: (2026)
Overcoming Data Inequality across Domains with Semi-Supervised Domain Generalization
di: Park, Jinha, et al.
Pubblicazione: (2024)
di: Park, Jinha, et al.
Pubblicazione: (2024)
Fast and Stable Diffusion Planning through Variational Adaptive Weighting
di: Qiu, Zhiying, et al.
Pubblicazione: (2025)
di: Qiu, Zhiying, et al.
Pubblicazione: (2025)
PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence
di: Xu, Yuanda, et al.
Pubblicazione: (2026)
di: Xu, Yuanda, et al.
Pubblicazione: (2026)
Continual Task Learning through Adaptive Policy Self-Composition
di: Hu, Shengchao, et al.
Pubblicazione: (2024)
di: Hu, Shengchao, et al.
Pubblicazione: (2024)
Action-Sufficient Goal Representations
di: Hyeon, Jinu, et al.
Pubblicazione: (2026)
di: Hyeon, Jinu, et al.
Pubblicazione: (2026)
Option-aware Temporally Abstracted Value for Offline Goal-Conditioned Reinforcement Learning
di: Ahn, Hongjoon, et al.
Pubblicazione: (2025)
di: Ahn, Hongjoon, et al.
Pubblicazione: (2025)
Listwise Reward Estimation for Offline Preference-based Reinforcement Learning
di: Choi, Heewoong, et al.
Pubblicazione: (2024)
di: Choi, Heewoong, et al.
Pubblicazione: (2024)
HDPO: Hybrid Distillation Policy Optimization via Privileged Self-Distillation
di: Ding, Ken
Pubblicazione: (2026)
di: Ding, Ken
Pubblicazione: (2026)
TIP: Token Importance in On-Policy Distillation
di: Xu, Yuanda, et al.
Pubblicazione: (2026)
di: Xu, Yuanda, et al.
Pubblicazione: (2026)
Online Policy Distillation with Decision-Attention
di: Yu, Xinqiang, et al.
Pubblicazione: (2024)
di: Yu, Xinqiang, et al.
Pubblicazione: (2024)
PromptKD: Distilling Student-Friendly Knowledge for Generative Language Models via Prompt Tuning
di: Kim, Gyeongman, et al.
Pubblicazione: (2024)
di: Kim, Gyeongman, et al.
Pubblicazione: (2024)
Stable Neural Stochastic Differential Equations in Analyzing Irregular Time Series Data
di: Oh, YongKyung, et al.
Pubblicazione: (2024)
di: Oh, YongKyung, et al.
Pubblicazione: (2024)
Exclusively Penalized Q-learning for Offline Reinforcement Learning
di: Yeom, Junghyuk, et al.
Pubblicazione: (2024)
di: Yeom, Junghyuk, et al.
Pubblicazione: (2024)
Fast Training of Sinusoidal Neural Fields via Scaling Initialization
di: Yeom, Taesun, et al.
Pubblicazione: (2024)
di: Yeom, Taesun, et al.
Pubblicazione: (2024)
Distilling Privileged Information for Dubins Traveling Salesman Problems with Neighborhoods
di: Shin, Min Kyu, et al.
Pubblicazione: (2024)
di: Shin, Min Kyu, et al.
Pubblicazione: (2024)
OPD+: Rethinking the Advantage Design for On-Policy Distillation
di: Zhao, Hanyang, et al.
Pubblicazione: (2026)
di: Zhao, Hanyang, et al.
Pubblicazione: (2026)
Trust-Region Behavior Blending for On-Policy Distillation
di: Plyusov, Daniil, et al.
Pubblicazione: (2026)
di: Plyusov, Daniil, et al.
Pubblicazione: (2026)
SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting
di: Zheng, Binbin, et al.
Pubblicazione: (2026)
di: Zheng, Binbin, et al.
Pubblicazione: (2026)
Cog-DRIFT: Exploration on Adaptively Reformulated Instances Enables Learning from Hard Reasoning Problems
di: Chen, Justin Chih-Yao, et al.
Pubblicazione: (2026)
di: Chen, Justin Chih-Yao, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Efficient Epistemic Uncertainty Estimation for Large Language Models via Knowledge Distillation
di: Park, Seonghyeon, et al.
Pubblicazione: (2026) -
Beyond Case Law: Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal QA
di: Chae, Kyubyung, et al.
Pubblicazione: (2026) -
Robust Domain Generalization under Divergent Marginal and Conditional Distributions
di: Yeom, Jewon, et al.
Pubblicazione: (2026) -
Patch-Level Kernel Alignment for Dense Self-Supervised Learning
di: Yeo, Juan, et al.
Pubblicazione: (2025) -
Towards Robust Real-World Multivariate Time Series Forecasting: A Unified Framework for Dependency, Asynchrony, and Missingness
di: Jang, Jinkwan, et al.
Pubblicazione: (2025)