Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Heyang, Yu, Xingrui, Bossens, David M., Tsang, Ivor W., Gu, Quanquan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Imitation from Diverse Behaviors: Wasserstein Quality Diversity Imitation Learning with Single-Step Archive Exploration
by: Yu, Xingrui, et al.
Published: (2024)
by: Yu, Xingrui, et al.
Published: (2024)
Diversifying Policy Behaviors with Extrinsic Behavioral Curiosity
by: Wan, Zhenglin, et al.
Published: (2024)
by: Wan, Zhenglin, et al.
Published: (2024)
Towards a Sharp Analysis of Offline Policy Learning for $f$-Divergence-Regularized Contextual Bandits
by: Zhao, Qingyue, et al.
Published: (2025)
by: Zhao, Qingyue, et al.
Published: (2025)
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
by: Zhao, Qingyue, et al.
Published: (2026)
by: Zhao, Qingyue, et al.
Published: (2026)
Mitigating Mismatch within Reference-based Preference Optimization
by: Yuan, Suqin, et al.
Published: (2026)
by: Yuan, Suqin, et al.
Published: (2026)
On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization
by: Ji, Kaixuan, et al.
Published: (2026)
by: Ji, Kaixuan, et al.
Published: (2026)
Near-Optimal Regret for KL-Regularized Multi-Armed Bandits
by: Ji, Kaixuan, et al.
Published: (2026)
by: Ji, Kaixuan, et al.
Published: (2026)
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations
by: Hoang, Huy, et al.
Published: (2025)
by: Hoang, Huy, et al.
Published: (2025)
Low Variance Off-policy Evaluation with State-based Importance Sampling
by: Bossens, David M., et al.
Published: (2022)
by: Bossens, David M., et al.
Published: (2022)
Uncertainty-Aware Reward-Free Exploration with General Function Approximation
by: Zhang, Junkai, et al.
Published: (2024)
by: Zhang, Junkai, et al.
Published: (2024)
Robust Lagrangian and Adversarial Policy Gradient for Robust Constrained Markov Decision Processes
by: Bossens, David M.
Published: (2023)
by: Bossens, David M.
Published: (2023)
Conceptual Belief-Informed Reinforcement Learning
by: Gu, Xingrui, et al.
Published: (2024)
by: Gu, Xingrui, et al.
Published: (2024)
How to Leverage Diverse Demonstrations in Offline Imitation Learning
by: Yue, Sheng, et al.
Published: (2024)
by: Yue, Sheng, et al.
Published: (2024)
Fast Direct: Query-Efficient Online Black-box Guidance for Diffusion-model Target Generation
by: Tan, Kim Yong, et al.
Published: (2025)
by: Tan, Kim Yong, et al.
Published: (2025)
Flow-Direct: Feedback-Efficient and Reusable Guidance for Flow Models via Non-Parametric Guidance Field
by: Tan, Kim Yong, et al.
Published: (2026)
by: Tan, Kim Yong, et al.
Published: (2026)
SPRINQL: Sub-optimal Demonstrations driven Offline Imitation Learning
by: Hoang, Huy, et al.
Published: (2024)
by: Hoang, Huy, et al.
Published: (2024)
Causal Imitation Learning under Expert-Observable and Expert-Unobservable Confounding
by: Shao, Daqian, et al.
Published: (2025)
by: Shao, Daqian, et al.
Published: (2025)
Variance-Dependent Regret Lower Bounds for Contextual Bandits
by: He, Jiafan, et al.
Published: (2025)
by: He, Jiafan, et al.
Published: (2025)
Imitation Learning from Suboptimal Demonstrations via Meta-Learning An Action Ranker
by: Fan, Jiangdong, et al.
Published: (2024)
by: Fan, Jiangdong, et al.
Published: (2024)
LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations
by: Ruoss, Anian, et al.
Published: (2024)
by: Ruoss, Anian, et al.
Published: (2024)
Sample-Efficient Expert Query Control in Active Imitation Learning via Conformal Prediction
by: Firouzkouhi, Arad, et al.
Published: (2025)
by: Firouzkouhi, Arad, et al.
Published: (2025)
Olaf-World: Orienting Latent Actions for Video World Modeling
by: Jiang, Yuxin, et al.
Published: (2026)
by: Jiang, Yuxin, et al.
Published: (2026)
IDIL: Imitation Learning of Intent-Driven Expert Behavior
by: Seo, Sangwon, et al.
Published: (2024)
by: Seo, Sangwon, et al.
Published: (2024)
Understanding Transferable Representation Learning and Zero-shot Transfer in CLIP
by: Chen, Zixiang, et al.
Published: (2023)
by: Chen, Zixiang, et al.
Published: (2023)
CoPS: Empowering LLM Agents with Provable Cross-Task Experience Sharing
by: Yang, Chen, et al.
Published: (2024)
by: Yang, Chen, et al.
Published: (2024)
Beyond Mimicry: Toward Lifelong Adaptability in Imitation Learning
by: Gavenski, Nathan, et al.
Published: (2026)
by: Gavenski, Nathan, et al.
Published: (2026)
An Integrated Imitation and Reinforcement Learning Methodology for Robust Agile Aircraft Control with Limited Pilot Demonstration Data
by: Sever, Gulay Goktas, et al.
Published: (2023)
by: Sever, Gulay Goktas, et al.
Published: (2023)
Efficient Imitation Without Demonstrations via Value-Penalized Auxiliary Control from Examples
by: Ablett, Trevor, et al.
Published: (2024)
by: Ablett, Trevor, et al.
Published: (2024)
Explorative Imitation Learning: A Path Signature Approach for Continuous Environments
by: Gavenski, Nathan, et al.
Published: (2024)
by: Gavenski, Nathan, et al.
Published: (2024)
Hierarchical Imitation Learning of Team Behavior from Heterogeneous Demonstrations
by: Seo, Sangwon, et al.
Published: (2025)
by: Seo, Sangwon, et al.
Published: (2025)
Global Convergence and Rich Feature Learning in $L$-Layer Infinite-Width Neural Networks under $μ$P Parametrization
by: Chen, Zixiang, et al.
Published: (2025)
by: Chen, Zixiang, et al.
Published: (2025)
Sharpness-Aware Black-Box Optimization
by: Ye, Feiyang, et al.
Published: (2024)
by: Ye, Feiyang, et al.
Published: (2024)
Advancing Analytic Class-Incremental Learning through Vision-Language Calibration
by: Zhao, Binyu, et al.
Published: (2026)
by: Zhao, Binyu, et al.
Published: (2026)
Distributional Multi-objective Black-box Optimization for Diffusion-model Inference-time Multi-Target Generation
by: Tan, Kim Yong, et al.
Published: (2025)
by: Tan, Kim Yong, et al.
Published: (2025)
Diffusion Language Models Can Perform Many Tasks with Scaling and Instruction-Finetuning
by: Ye, Jiasheng, et al.
Published: (2023)
by: Ye, Jiasheng, et al.
Published: (2023)
Imitation Learning for Multi-turn LM Agents via On-policy Expert Corrections
by: Lauffer, Niklas, et al.
Published: (2025)
by: Lauffer, Niklas, et al.
Published: (2025)
Coupled Distributional Random Expert Distillation for World Model Online Imitation Learning
by: Li, Shangzhe, et al.
Published: (2025)
by: Li, Shangzhe, et al.
Published: (2025)
Dual-Balancing for Multi-Task Learning
by: Lin, Baijiong, et al.
Published: (2023)
by: Lin, Baijiong, et al.
Published: (2023)
PAIL: Performance based Adversarial Imitation Learning Engine for Carbon Neutral Optimization
by: Ye, Yuyang, et al.
Published: (2024)
by: Ye, Yuyang, et al.
Published: (2024)
Transductive Reward Inference on Graph
by: Qu, Bohao, et al.
Published: (2024)
by: Qu, Bohao, et al.
Published: (2024)
Similar Items
-
Imitation from Diverse Behaviors: Wasserstein Quality Diversity Imitation Learning with Single-Step Archive Exploration
by: Yu, Xingrui, et al.
Published: (2024) -
Diversifying Policy Behaviors with Extrinsic Behavioral Curiosity
by: Wan, Zhenglin, et al.
Published: (2024) -
Towards a Sharp Analysis of Offline Policy Learning for $f$-Divergence-Regularized Contextual Bandits
by: Zhao, Qingyue, et al.
Published: (2025) -
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
by: Zhao, Qingyue, et al.
Published: (2026) -
Mitigating Mismatch within Reference-based Preference Optimization
by: Yuan, Suqin, et al.
Published: (2026)