Reward-free World Models for Online Imitation Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Shangzhe, Huang, Zhiao, Su, Hao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Coupled Distributional Random Expert Distillation for World Model Online Imitation Learning
by: Li, Shangzhe, et al.
Published: (2025)
by: Li, Shangzhe, et al.
Published: (2025)
A Recipe for Efficient Sim-to-Real Transfer in Manipulation with Online Imitation-Pretrained World Models
by: Wang, Yilin, et al.
Published: (2025)
by: Wang, Yilin, et al.
Published: (2025)
Near-Optimal Second-Order Guarantees for Model-Based Adversarial Imitation Learning
by: Li, Shangzhe, et al.
Published: (2025)
by: Li, Shangzhe, et al.
Published: (2025)
Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning
by: Li, Shangzhe, et al.
Published: (2026)
by: Li, Shangzhe, et al.
Published: (2026)
Imitation from Observations with Trajectory-Level Generative Embeddings
by: Qu, Yongtao, et al.
Published: (2026)
by: Qu, Yongtao, et al.
Published: (2026)
Provably Efficient Offline-to-Online Value Adaptation with General Function Approximation
by: Li, Shangzhe, et al.
Published: (2026)
by: Li, Shangzhe, et al.
Published: (2026)
Augmenting Offline Reinforcement Learning with State-only Interactions
by: Li, Shangzhe, et al.
Published: (2024)
by: Li, Shangzhe, et al.
Published: (2024)
Stealthy Imitation: Reward-guided Environment-free Policy Stealing
by: Zhuang, Zhixiong, et al.
Published: (2024)
by: Zhuang, Zhixiong, et al.
Published: (2024)
Overcoming Knowledge Barriers: Online Imitation Learning from Visual Observation with Pretrained World Models
by: Zhang, Xingyuan, et al.
Published: (2024)
by: Zhang, Xingyuan, et al.
Published: (2024)
Uncovering Capabilities of Model Pruning in Graph Contrastive Learning
by: Wu, Junran, et al.
Published: (2024)
by: Wu, Junran, et al.
Published: (2024)
Efficient Imitation Learning with Conservative World Models
by: Kolev, Victor, et al.
Published: (2024)
by: Kolev, Victor, et al.
Published: (2024)
Expert Proximity as Surrogate Rewards for Single Demonstration Imitation Learning
by: Chiang, Chia-Cheng, et al.
Published: (2024)
by: Chiang, Chia-Cheng, et al.
Published: (2024)
PAGAR: Taming Reward Misalignment in Inverse Reinforcement Learning-Based Imitation Learning with Protagonist Antagonist Guided Adversarial Reward
by: Zhou, Weichao, et al.
Published: (2023)
by: Zhou, Weichao, et al.
Published: (2023)
Diffusion-Reward Adversarial Imitation Learning
by: Lai, Chun-Mao, et al.
Published: (2024)
by: Lai, Chun-Mao, et al.
Published: (2024)
DITTO: Offline Imitation Learning with World Models
by: DeMoss, Branton, et al.
Published: (2023)
by: DeMoss, Branton, et al.
Published: (2023)
Chain-of-Thought Predictive Control
by: Jia, Zhiwei, et al.
Published: (2023)
by: Jia, Zhiwei, et al.
Published: (2023)
Quantile Q-Learning: Revisiting Offline Extreme Q-Learning with Quantile Regression
by: Gao, Xinming, et al.
Published: (2025)
by: Gao, Xinming, et al.
Published: (2025)
From Novelty to Imitation: Self-Distilled Rewards for Offline Reinforcement Learning
by: Chaudhary, Gaurav, et al.
Published: (2025)
by: Chaudhary, Gaurav, et al.
Published: (2025)
Multi-Stage Manipulation with Demonstration-Augmented Reward, Policy, and World Model Learning
by: Escoriza, Adrià López, et al.
Published: (2025)
by: Escoriza, Adrià López, et al.
Published: (2025)
Fine-Tuning Language Models with Reward Learning on Policy
by: Lang, Hao, et al.
Published: (2024)
by: Lang, Hao, et al.
Published: (2024)
Online Adaptation for Enhancing Imitation Learning Policies
by: Malato, Federico, et al.
Published: (2024)
by: Malato, Federico, et al.
Published: (2024)
Beyond Imitation: Recovering Dense Rewards from Demonstrations
by: Li, Jiangnan, et al.
Published: (2025)
by: Li, Jiangnan, et al.
Published: (2025)
Rethinking Adversarial Inverse Reinforcement Learning: Policy Imitation, Transferable Reward Recovery and Algebraic Equilibrium Proof
by: Zhang, Yangchun, et al.
Published: (2024)
by: Zhang, Yangchun, et al.
Published: (2024)
Extrinsicaly Rewarded Soft Q Imitation Learning with Discriminator
by: Furuyama, Ryoma, et al.
Published: (2024)
by: Furuyama, Ryoma, et al.
Published: (2024)
Offline Imitation Learning with Model-based Reverse Augmentation
by: Shao, Jie-Jing, et al.
Published: (2024)
by: Shao, Jie-Jing, et al.
Published: (2024)
PROF: An LLM-based Reward Code Preference Optimization Framework for Offline Imitation Learning
by: Sun, Shengjie, et al.
Published: (2025)
by: Sun, Shengjie, et al.
Published: (2025)
LUMOS: Language-Conditioned Imitation Learning with World Models
by: Nematollahi, Iman, et al.
Published: (2025)
by: Nematollahi, Iman, et al.
Published: (2025)
Self-evolved Imitation Learning in Simulated World
by: Ye, Yifan, et al.
Published: (2025)
by: Ye, Yifan, et al.
Published: (2025)
Curriculum Learning and Imitation Learning for Model-free Control on Financial Time-series
by: Koh, Woosung, et al.
Published: (2023)
by: Koh, Woosung, et al.
Published: (2023)
Provably Efficient Online RLHF with One-Pass Reward Modeling
by: Li, Long-Fei, et al.
Published: (2025)
by: Li, Long-Fei, et al.
Published: (2025)
OLLIE: Imitation Learning from Offline Pretraining to Online Finetuning
by: Yue, Sheng, et al.
Published: (2024)
by: Yue, Sheng, et al.
Published: (2024)
Off-Dynamics Reinforcement Learning via Domain Adaptation and Reward Augmented Imitation
by: Guo, Yihong, et al.
Published: (2024)
by: Guo, Yihong, et al.
Published: (2024)
DrS: Learning Reusable Dense Rewards for Multi-Stage Tasks
by: Mu, Tongzhou, et al.
Published: (2024)
by: Mu, Tongzhou, et al.
Published: (2024)
A Note on Hybrid Online Reinforcement and Imitation Learning for LLMs: Formulations and Algorithms
by: Li, Yingru, et al.
Published: (2025)
by: Li, Yingru, et al.
Published: (2025)
Riemannian Projection-free Online Learning
by: Hu, Zihao, et al.
Published: (2023)
by: Hu, Zihao, et al.
Published: (2023)
Inverse Contextual Bandits without Rewards: Learning from a Non-Stationary Learner via Suffix Imitation
by: Kong, Yuqi, et al.
Published: (2026)
by: Kong, Yuqi, et al.
Published: (2026)
Online Imitation Learning for Manipulation via Decaying Relative Correction through Teleoperation
by: Pan, Cheng, et al.
Published: (2025)
by: Pan, Cheng, et al.
Published: (2025)
Trajectory World Models for Heterogeneous Environments
by: Yin, Shaofeng, et al.
Published: (2025)
by: Yin, Shaofeng, et al.
Published: (2025)
Reward-Free Curricula for Training Robust World Models
by: Rigter, Marc, et al.
Published: (2023)
by: Rigter, Marc, et al.
Published: (2023)
Learning for Long-Horizon Planning via Neuro-Symbolic Abductive Imitation
by: Shao, Jie-Jing, et al.
Published: (2024)
by: Shao, Jie-Jing, et al.
Published: (2024)
Similar Items
-
Coupled Distributional Random Expert Distillation for World Model Online Imitation Learning
by: Li, Shangzhe, et al.
Published: (2025) -
A Recipe for Efficient Sim-to-Real Transfer in Manipulation with Online Imitation-Pretrained World Models
by: Wang, Yilin, et al.
Published: (2025) -
Near-Optimal Second-Order Guarantees for Model-Based Adversarial Imitation Learning
by: Li, Shangzhe, et al.
Published: (2025) -
Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning
by: Li, Shangzhe, et al.
Published: (2026) -
Imitation from Observations with Trajectory-Level Generative Embeddings
by: Qu, Yongtao, et al.
Published: (2026)