Adversarial Dual On-Policy Distillation from Expressive Teacher

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wan, Zhenglin, Wu, Jingxuan, Yu, Xingrui, Zhang, Chubin, Lei, Mingcong, An, Bo, Tsang, Ivor W., You, Yang
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914622707073024
author Wan, Zhenglin
Wu, Jingxuan
Yu, Xingrui
Zhang, Chubin
Lei, Mingcong
An, Bo
Tsang, Ivor W.
You, Yang
author_facet Wan, Zhenglin
Wu, Jingxuan
Yu, Xingrui
Zhang, Chubin
Lei, Mingcong
An, Bo
Tsang, Ivor W.
You, Yang
contents Learning from demonstrations in embodied control is often cast as behavioral cloning, and recent diffusion or flow-matching policies improve this paradigm by modeling multi-modal expert actions. Yet these methods remain offline supervised learners: the policy is trained only on expert states and receives no corrective signal on the states it actually visits. On-policy distillation (OPD) offers a natural remedy, but standard OPD assumes a strong fixed teacher, which is unavailable in demonstration-only control. We propose \textbf{FA-OPD}, an \emph{adversarial dual on-policy distillation} method in which a Flow Matching (FM) teacher is learned from demonstrations and co-trained with a lightweight MLP student. The teacher provides two complementary signals on student rollouts. The reward channel learns an expert-likeness objective over state-action pairs and drives online exploration through long-horizon policy optimization. The action channel supplies dense local targets at student-visited states, stabilizing exploitation. FA-OPD couples them so that reward distillation enables generalization beyond point-wise demonstrations, while action distillation keeps exploration anchored near expert-like behavior. Across six robot navigation, manipulation, and locomotion benchmarks, FA-OPD beats strong baselines and shows much stronger robustness under noisy or limited demonstrations. Source code: https://github.com/vanzll/FA-OPD.
format Preprint
id arxiv_https___arxiv_org_abs_2605_27095
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Adversarial Dual On-Policy Distillation from Expressive Teacher
Wan, Zhenglin
Wu, Jingxuan
Yu, Xingrui
Zhang, Chubin
Lei, Mingcong
An, Bo
Tsang, Ivor W.
You, Yang
Machine Learning
Learning from demonstrations in embodied control is often cast as behavioral cloning, and recent diffusion or flow-matching policies improve this paradigm by modeling multi-modal expert actions. Yet these methods remain offline supervised learners: the policy is trained only on expert states and receives no corrective signal on the states it actually visits. On-policy distillation (OPD) offers a natural remedy, but standard OPD assumes a strong fixed teacher, which is unavailable in demonstration-only control. We propose \textbf{FA-OPD}, an \emph{adversarial dual on-policy distillation} method in which a Flow Matching (FM) teacher is learned from demonstrations and co-trained with a lightweight MLP student. The teacher provides two complementary signals on student rollouts. The reward channel learns an expert-likeness objective over state-action pairs and drives online exploration through long-horizon policy optimization. The action channel supplies dense local targets at student-visited states, stabilizing exploitation. FA-OPD couples them so that reward distillation enables generalization beyond point-wise demonstrations, while action distillation keeps exploration anchored near expert-like behavior. Across six robot navigation, manipulation, and locomotion benchmarks, FA-OPD beats strong baselines and shows much stronger robustness under noisy or limited demonstrations. Source code: https://github.com/vanzll/FA-OPD.
title Adversarial Dual On-Policy Distillation from Expressive Teacher
topic Machine Learning
url https://arxiv.org/abs/2605.27095