Adversarial Dual On-Policy Distillation from Expressive Teacher
Fuente:
arXiv
Saved in:
| Main Authors: | Wan, Zhenglin, Wu, Jingxuan, Yu, Xingrui, Zhang, Chubin, Lei, Mingcong, An, Bo, Tsang, Ivor W., You, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning
by: Wan, Zhenglin, et al.
Published: (2025)
by: Wan, Zhenglin, et al.
Published: (2025)
Evolving Diffusion and Flow Matching Policies for Online Reinforcement Learning
by: Zhang, Chubin, et al.
Published: (2025)
by: Zhang, Chubin, et al.
Published: (2025)
HBVLA: Pushing 1-Bit Post-Training Quantization for Vision-Language-Action Models
by: Yan, Xin, et al.
Published: (2026)
by: Yan, Xin, et al.
Published: (2026)
Letting Trajectories Spread: Quality-Preserving Control for Diverse Flow Matching
by: Wu, Jingxuan, et al.
Published: (2025)
by: Wu, Jingxuan, et al.
Published: (2025)
Imitation from Diverse Behaviors: Wasserstein Quality Diversity Imitation Learning with Single-Step Archive Exploration
by: Yu, Xingrui, et al.
Published: (2024)
by: Yu, Xingrui, et al.
Published: (2024)
Diversifying Policy Behaviors with Extrinsic Behavioral Curiosity
by: Wan, Zhenglin, et al.
Published: (2024)
by: Wan, Zhenglin, et al.
Published: (2024)
Time-Annealed Perturbation Sampling: Diverse Generation for Diffusion Language Models
by: Wu, Jingxuan, et al.
Published: (2026)
by: Wu, Jingxuan, et al.
Published: (2026)
Advancing Analytic Class-Incremental Learning through Vision-Language Calibration
by: Zhao, Binyu, et al.
Published: (2026)
by: Zhao, Binyu, et al.
Published: (2026)
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration
by: Zhao, Heyang, et al.
Published: (2025)
by: Zhao, Heyang, et al.
Published: (2025)
Mitigating Mismatch within Reference-based Preference Optimization
by: Yuan, Suqin, et al.
Published: (2026)
by: Yuan, Suqin, et al.
Published: (2026)
Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe
by: Hou, Wenjin, et al.
Published: (2026)
by: Hou, Wenjin, et al.
Published: (2026)
IKNO: Infinite-order Kernel Neural Operators
by: Zhu, Pengyuan, et al.
Published: (2026)
by: Zhu, Pengyuan, et al.
Published: (2026)
Multi-Modal Dataset Distillation in the Wild
by: Dang, Zhuohang, et al.
Published: (2025)
by: Dang, Zhuohang, et al.
Published: (2025)
A Teacher-Free Graph Knowledge Distillation Framework with Dual Self-Distillation
by: Wu, Lirong, et al.
Published: (2024)
by: Wu, Lirong, et al.
Published: (2024)
Policy Dispersion in Non-Markovian Environment
by: Qu, Bohao, et al.
Published: (2023)
by: Qu, Bohao, et al.
Published: (2023)
A First-Order Multi-Gradient Algorithm for Multi-Objective Bi-Level Optimization
by: Ye, Feiyang, et al.
Published: (2024)
by: Ye, Feiyang, et al.
Published: (2024)
Dual-Balancing for Multi-Task Learning
by: Lin, Baijiong, et al.
Published: (2023)
by: Lin, Baijiong, et al.
Published: (2023)
Training-Free Dataset Pruning for Instance Segmentation
by: Dai, Yalun, et al.
Published: (2025)
by: Dai, Yalun, et al.
Published: (2025)
Primal-Dual Policy Optimization for Linear CMDPs with Adversarial Losses
by: Yu, Kihyun, et al.
Published: (2026)
by: Yu, Kihyun, et al.
Published: (2026)
Learning ORDER-Aware Multimodal Representations for Composite Materials Design
by: Li, Xinyao, et al.
Published: (2026)
by: Li, Xinyao, et al.
Published: (2026)
The Propagation Field: A Geometric Substrate Theory of Deep Learning
by: Gu, Xingrui
Published: (2026)
by: Gu, Xingrui
Published: (2026)
Covariance-Adaptive Sequential Black-box Optimization for Diffusion Targeted Generation
by: Lyu, Yueming, et al.
Published: (2024)
by: Lyu, Yueming, et al.
Published: (2024)
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
by: Yang, Wenkai, et al.
Published: (2026)
by: Yang, Wenkai, et al.
Published: (2026)
Continual Policy Distillation from Distributed Reinforcement Learning Teachers
by: Li, Yuxuan, et al.
Published: (2026)
by: Li, Yuxuan, et al.
Published: (2026)
Diversified Batch Selection for Training Acceleration
by: Hong, Feng, et al.
Published: (2024)
by: Hong, Feng, et al.
Published: (2024)
Towards Harmless Rawlsian Fairness Regardless of Demographic Prior
by: Wang, Xuanqian, et al.
Published: (2024)
by: Wang, Xuanqian, et al.
Published: (2024)
Uncertainty-Gated Generative Modeling
by: Gu, Xingrui, et al.
Published: (2026)
by: Gu, Xingrui, et al.
Published: (2026)
ANO: A Principled Approach to Robust Policy Optimization
by: Zhang, Yiheng, et al.
Published: (2026)
by: Zhang, Yiheng, et al.
Published: (2026)
Dual-Forward Path Teacher Knowledge Distillation: Bridging the Capacity Gap Between Teacher and Student
by: Li, Tong, et al.
Published: (2025)
by: Li, Tong, et al.
Published: (2025)
Lang-PINN: From Language to Physics-Informed Neural Networks via a Multi-Agent Framework
by: He, Xin, et al.
Published: (2025)
by: He, Xin, et al.
Published: (2025)
Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning
by: Yang, Xuewei, et al.
Published: (2026)
by: Yang, Xuewei, et al.
Published: (2026)
Toward Fair Graph Neural Networks Via Dual-Teacher Knowledge Distillation
by: Li, Chengyu, et al.
Published: (2024)
by: Li, Chengyu, et al.
Published: (2024)
Sharpness-Aware Black-Box Optimization
by: Ye, Feiyang, et al.
Published: (2024)
by: Ye, Feiyang, et al.
Published: (2024)
PROUD: PaRetO-gUided Diffusion Model for Multi-objective Generation
by: Yao, Yinghua, et al.
Published: (2024)
by: Yao, Yinghua, et al.
Published: (2024)
Exploring Hierarchical Molecular Graph Representation in Multimodal LLMs
by: Hu, Chengxin, et al.
Published: (2024)
by: Hu, Chengxin, et al.
Published: (2024)
Annealing Self-Distillation Rectification Improves Adversarial Training
by: Wu, Yu-Yu, et al.
Published: (2023)
by: Wu, Yu-Yu, et al.
Published: (2023)
Cross-Context Backdoor Attacks against Graph Prompt Learning
by: Lyu, Xiaoting, et al.
Published: (2024)
by: Lyu, Xiaoting, et al.
Published: (2024)
On Expressivity of Height in Neural Networks
by: Fan, Feng-Lei, et al.
Published: (2023)
by: Fan, Feng-Lei, et al.
Published: (2023)
Disentangling Structured Components: Towards Adaptive, Interpretable and Scalable Time Series Forecasting
by: Deng, Jinliang, et al.
Published: (2023)
by: Deng, Jinliang, et al.
Published: (2023)
Alpha and Prejudice: Improving $α$-sized Worst-case Fairness via Intrinsic Reweighting
by: Li, Jing, et al.
Published: (2024)
by: Li, Jing, et al.
Published: (2024)
Similar Items
-
FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning
by: Wan, Zhenglin, et al.
Published: (2025) -
Evolving Diffusion and Flow Matching Policies for Online Reinforcement Learning
by: Zhang, Chubin, et al.
Published: (2025) -
HBVLA: Pushing 1-Bit Post-Training Quantization for Vision-Language-Action Models
by: Yan, Xin, et al.
Published: (2026) -
Letting Trajectories Spread: Quality-Preserving Control for Diverse Flow Matching
by: Wu, Jingxuan, et al.
Published: (2025) -
Imitation from Diverse Behaviors: Wasserstein Quality Diversity Imitation Learning with Single-Step Archive Exploration
by: Yu, Xingrui, et al.
Published: (2024)