Efficient Imitation Learning with Conservative World Models
Fuente:
arXiv
Saved in:
| Main Authors: | Kolev, Victor, Rafailov, Rafael, Hatch, Kyle, Wu, Jiajun, Finn, Chelsea |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MOTO: Offline Pre-training to Online Fine-tuning for Model-based Robot Learning
by: Rafailov, Rafael, et al.
Published: (2024)
by: Rafailov, Rafael, et al.
Published: (2024)
D5RL: Diverse Datasets for Data-Driven Deep Reinforcement Learning
by: Rafailov, Rafael, et al.
Published: (2024)
by: Rafailov, Rafael, et al.
Published: (2024)
From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function
by: Rafailov, Rafael, et al.
Published: (2024)
by: Rafailov, Rafael, et al.
Published: (2024)
Disentangling Length from Quality in Direct Preference Optimization
by: Park, Ryan, et al.
Published: (2024)
by: Park, Ryan, et al.
Published: (2024)
Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
by: Xiang, Violet, et al.
Published: (2025)
by: Xiang, Violet, et al.
Published: (2025)
Aligning Modalities in Vision Large Language Models via Preference Fine-tuning
by: Zhou, Yiyang, et al.
Published: (2024)
by: Zhou, Yiyang, et al.
Published: (2024)
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
by: Rafailov, Rafael, et al.
Published: (2023)
by: Rafailov, Rafael, et al.
Published: (2023)
Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents
by: Putta, Pranav, et al.
Published: (2024)
by: Putta, Pranav, et al.
Published: (2024)
Contrastive Preference Learning: Learning from Human Feedback without RL
by: Hejna, Joey, et al.
Published: (2023)
by: Hejna, Joey, et al.
Published: (2023)
Reinforcement Learning via Implicit Imitation Guidance
by: Dong, Perry, et al.
Published: (2025)
by: Dong, Perry, et al.
Published: (2025)
Tripod: Three Complementary Inductive Biases for Disentangled Representation Learning
by: Hsu, Kyle, et al.
Published: (2024)
by: Hsu, Kyle, et al.
Published: (2024)
Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
by: Rafailov, Rafael, et al.
Published: (2024)
by: Rafailov, Rafael, et al.
Published: (2024)
Generative Reward Models
by: Mahan, Dakota, et al.
Published: (2024)
by: Mahan, Dakota, et al.
Published: (2024)
Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data
by: Tajwar, Fahim, et al.
Published: (2024)
by: Tajwar, Fahim, et al.
Published: (2024)
HumanPlus: Humanoid Shadowing and Imitation from Humans
by: Fu, Zipeng, et al.
Published: (2024)
by: Fu, Zipeng, et al.
Published: (2024)
Dual-Force: Enhanced Offline Diversity Maximization under Imitation Constraints
by: Kolev, Pavel, et al.
Published: (2025)
by: Kolev, Pavel, et al.
Published: (2025)
Offline Diversity Maximization Under Imitation Constraints
by: Vlastelica, Marin, et al.
Published: (2023)
by: Vlastelica, Marin, et al.
Published: (2023)
Conservative Prediction via Data-Driven Confidence Minimization
by: Choi, Caroline, et al.
Published: (2023)
by: Choi, Caroline, et al.
Published: (2023)
FAST: Efficient Action Tokenization for Vision-Language-Action Models
by: Pertsch, Karl, et al.
Published: (2025)
by: Pertsch, Karl, et al.
Published: (2025)
Reward-free World Models for Online Imitation Learning
by: Li, Shangzhe, et al.
Published: (2024)
by: Li, Shangzhe, et al.
Published: (2024)
DITTO: Offline Imitation Learning with World Models
by: DeMoss, Branton, et al.
Published: (2023)
by: DeMoss, Branton, et al.
Published: (2023)
Helpful DoggyBot: Open-World Object Fetching using Legged Robots and Vision-Language Models
by: Wu, Qi, et al.
Published: (2024)
by: Wu, Qi, et al.
Published: (2024)
Feedback Descent: Open-Ended Text Optimization via Pairwise Comparison
by: Lee, Yoonho, et al.
Published: (2025)
by: Lee, Yoonho, et al.
Published: (2025)
EXPO: Stable Reinforcement Learning with Expressive Policies
by: Dong, Perry, et al.
Published: (2025)
by: Dong, Perry, et al.
Published: (2025)
Grounding by Trying: LLMs with Reinforcement Learning-Enhanced Retrieval
by: Hsu, Sheryl, et al.
Published: (2024)
by: Hsu, Sheryl, et al.
Published: (2024)
Self-Guided Masked Autoencoders for Domain-Agnostic Self-Supervised Learning
by: Xie, Johnathan, et al.
Published: (2024)
by: Xie, Johnathan, et al.
Published: (2024)
Universal Neural Functionals
by: Zhou, Allan, et al.
Published: (2024)
by: Zhou, Allan, et al.
Published: (2024)
Learning Long-Context Diffusion Policies via Past-Token Prediction
by: Torne, Marcel, et al.
Published: (2025)
by: Torne, Marcel, et al.
Published: (2025)
Efficient Data Collection for Robotic Manipulation via Compositional Generalization
by: Gao, Jensen, et al.
Published: (2024)
by: Gao, Jensen, et al.
Published: (2024)
Collapse or Thrive? Perils and Promises of Synthetic Data in a Self-Generating World
by: Kazdan, Joshua, et al.
Published: (2024)
by: Kazdan, Joshua, et al.
Published: (2024)
World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry
by: Liu, Yuejiang, et al.
Published: (2026)
by: Liu, Yuejiang, et al.
Published: (2026)
Posterior Behavioral Cloning: Pretraining BC Policies for Efficient RL Finetuning
by: Wagenmaker, Andrew, et al.
Published: (2025)
by: Wagenmaker, Andrew, et al.
Published: (2025)
LUMOS: Language-Conditioned Imitation Learning with World Models
by: Nematollahi, Iman, et al.
Published: (2025)
by: Nematollahi, Iman, et al.
Published: (2025)
Self-evolved Imitation Learning in Simulated World
by: Ye, Yifan, et al.
Published: (2025)
by: Ye, Yifan, et al.
Published: (2025)
Efficient Active Imitation Learning with Random Network Distillation
by: Biré, Emilien, et al.
Published: (2024)
by: Biré, Emilien, et al.
Published: (2024)
Efficient Offline Reinforcement Learning: First Imitate, then Improve
by: Jelley, Adam, et al.
Published: (2024)
by: Jelley, Adam, et al.
Published: (2024)
Overcoming Knowledge Barriers: Online Imitation Learning from Visual Observation with Pretrained World Models
by: Zhang, Xingyuan, et al.
Published: (2024)
by: Zhang, Xingyuan, et al.
Published: (2024)
Coupled Distributional Random Expert Distillation for World Model Online Imitation Learning
by: Li, Shangzhe, et al.
Published: (2025)
by: Li, Shangzhe, et al.
Published: (2025)
Evaluating Real-World Robot Manipulation Policies in Simulation
by: Li, Xuanlin, et al.
Published: (2024)
by: Li, Xuanlin, et al.
Published: (2024)
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
by: Kim, Moo Jin, et al.
Published: (2025)
by: Kim, Moo Jin, et al.
Published: (2025)
Similar Items
-
MOTO: Offline Pre-training to Online Fine-tuning for Model-based Robot Learning
by: Rafailov, Rafael, et al.
Published: (2024) -
D5RL: Diverse Datasets for Data-Driven Deep Reinforcement Learning
by: Rafailov, Rafael, et al.
Published: (2024) -
From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function
by: Rafailov, Rafael, et al.
Published: (2024) -
Disentangling Length from Quality in Direct Preference Optimization
by: Park, Ryan, et al.
Published: (2024) -
Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
by: Xiang, Violet, et al.
Published: (2025)