VILP: Imitation Learning with Latent Video Planning
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Zhengtong, Qiu, Qiang, She, Yu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Visual Feature-Based World Models via Residual Latent Action
by: Zhang, Xinyu, et al.
Published: (2026)
by: Zhang, Xinyu, et al.
Published: (2026)
Lifelong Imitation Learning with Multimodal Latent Replay and Incremental Adjustment
by: Yu, Fanqi, et al.
Published: (2026)
by: Yu, Fanqi, et al.
Published: (2026)
UNIC: Learning Unified Multimodal Extrinsic Contact Estimation
by: Xu, Zhengtong, et al.
Published: (2026)
by: Xu, Zhengtong, et al.
Published: (2026)
EgoMimic: Scaling Imitation Learning via Egocentric Video
by: Kareer, Simar, et al.
Published: (2024)
by: Kareer, Simar, et al.
Published: (2024)
PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and Planning
by: Villar-Corrales, Angel, et al.
Published: (2025)
by: Villar-Corrales, Angel, et al.
Published: (2025)
Contrastive Imitation Learning for Language-guided Multi-Task Robotic Manipulation
by: Ma, Teli, et al.
Published: (2024)
by: Ma, Teli, et al.
Published: (2024)
RoboMirror: Understand Before You Imitate for Video to Humanoid Locomotion
by: Li, Zhe, et al.
Published: (2025)
by: Li, Zhe, et al.
Published: (2025)
Slot-Level Robotic Placement via Visual Imitation from Single Human Video
by: Shan, Dandan, et al.
Published: (2025)
by: Shan, Dandan, et al.
Published: (2025)
One-shot Video Imitation via Parameterized Symbolic Abstraction Graphs
by: Wang, Jianren, et al.
Published: (2024)
by: Wang, Jianren, et al.
Published: (2024)
UniSkill: Imitating Human Videos via Cross-Embodiment Skill Representations
by: Kim, Hanjung, et al.
Published: (2025)
by: Kim, Hanjung, et al.
Published: (2025)
One-Shot Dual-Arm Imitation Learning
by: Wang, Yilong, et al.
Published: (2025)
by: Wang, Yilong, et al.
Published: (2025)
CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning
by: Yang, Jiange, et al.
Published: (2025)
by: Yang, Jiange, et al.
Published: (2025)
Uni-World VLA: Interleaved World Modeling and Planning for Autonomous Driving
by: Liu, Qiqi, et al.
Published: (2026)
by: Liu, Qiqi, et al.
Published: (2026)
LLM-enhanced Scene Graph Learning for Household Rearrangement
by: Li, Wenhao, et al.
Published: (2024)
by: Li, Wenhao, et al.
Published: (2024)
Towards Fusing Point Cloud and Visual Representations for Imitation Learning
by: Donat, Atalay, et al.
Published: (2025)
by: Donat, Atalay, et al.
Published: (2025)
Failure Identification in Imitation Learning Via Statistical and Semantic Filtering
by: Rolland, Quentin, et al.
Published: (2026)
by: Rolland, Quentin, et al.
Published: (2026)
LIDEA: Human-to-Robot Imitation Learning via Implicit Feature Distillation and Explicit Geometry Alignment
by: Xu, Yifu, et al.
Published: (2026)
by: Xu, Yifu, et al.
Published: (2026)
CoIRL-AD: Collaborative-Competitive Imitation-Reinforcement Learning in Latent World Models for Autonomous Driving
by: Zheng, Xiaoji, et al.
Published: (2025)
by: Zheng, Xiaoji, et al.
Published: (2025)
Stem-OB: Generalizable Visual Imitation Learning with Stem-Like Convergent Observation through Diffusion Inversion
by: Hu, Kaizhe, et al.
Published: (2024)
by: Hu, Kaizhe, et al.
Published: (2024)
FUNCTO: Function-Centric One-Shot Imitation Learning for Tool Manipulation
by: Tang, Chao, et al.
Published: (2025)
by: Tang, Chao, et al.
Published: (2025)
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
by: Zhang, Chubin, et al.
Published: (2026)
by: Zhang, Chubin, et al.
Published: (2026)
Imitating What Works: Simulation-Filtered Modular Policy Learning from Human Videos
by: Zhai, Albert J., et al.
Published: (2026)
by: Zhai, Albert J., et al.
Published: (2026)
Learning Sidewalk Autopilot from Multi-Scale Imitation with Corrective Behavior Expansion
by: He, Honglin, et al.
Published: (2026)
by: He, Honglin, et al.
Published: (2026)
Dream to Manipulate: Compositional World Models Empowering Robot Imitation Learning with Imagination
by: Barcellona, Leonardo, et al.
Published: (2024)
by: Barcellona, Leonardo, et al.
Published: (2024)
DITTO: Demonstration Imitation by Trajectory Transformation
by: Heppert, Nick, et al.
Published: (2024)
by: Heppert, Nick, et al.
Published: (2024)
Beyond Imitation: Learning Safe End-to-End Autonomous Driving from Hard Negatives
by: Wang, Junli, et al.
Published: (2026)
by: Wang, Junli, et al.
Published: (2026)
Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations
by: Patel, Shivansh, et al.
Published: (2025)
by: Patel, Shivansh, et al.
Published: (2025)
WorldRFT: Latent World Model Planning with Reinforcement Fine-Tuning for Autonomous Driving
by: Yang, Pengxuan, et al.
Published: (2025)
by: Yang, Pengxuan, et al.
Published: (2025)
Visual Imitation Enables Contextual Humanoid Control
by: Allshire, Arthur, et al.
Published: (2025)
by: Allshire, Arthur, et al.
Published: (2025)
NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation
by: Zhang, Jiazhao, et al.
Published: (2024)
by: Zhang, Jiazhao, et al.
Published: (2024)
Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving
by: Wang, Linbo, et al.
Published: (2026)
by: Wang, Linbo, et al.
Published: (2026)
Efficient Robotic Policy Learning via Latent Space Backward Planning
by: Liu, Dongxiu, et al.
Published: (2025)
by: Liu, Dongxiu, et al.
Published: (2025)
DIVER: Reinforced Diffusion Breaks Imitation Bottlenecks in End-to-End Autonomous Driving
by: Song, Ziying, et al.
Published: (2025)
by: Song, Ziying, et al.
Published: (2025)
Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation
by: Lu, Jinghui, et al.
Published: (2026)
by: Lu, Jinghui, et al.
Published: (2026)
Robust Instant Policy: Leveraging Student's t-Regression Model for Robust In-context Imitation Learning of Robot Manipulation
by: Oh, Hanbit, et al.
Published: (2025)
by: Oh, Hanbit, et al.
Published: (2025)
UAV-Flow Colosseo: A Real-World Benchmark for Flying-on-a-Word UAV Imitation Learning
by: Wang, Xiangyu, et al.
Published: (2025)
by: Wang, Xiangyu, et al.
Published: (2025)
SPIRE: Synergistic Planning, Imitation, and Reinforcement Learning for Long-Horizon Manipulation
by: Zhou, Zihan, et al.
Published: (2024)
by: Zhou, Zihan, et al.
Published: (2024)
Solving Motion Planning Tasks with a Scalable Generative Model
by: Hu, Yihan, et al.
Published: (2024)
by: Hu, Yihan, et al.
Published: (2024)
GenH2R: Learning Generalizable Human-to-Robot Handover via Scalable Simulation, Demonstration, and Imitation
by: Wang, Zifan, et al.
Published: (2024)
by: Wang, Zifan, et al.
Published: (2024)
LatentBKI: Open-Dictionary Continuous Mapping in Visual-Language Latent Spaces with Quantifiable Uncertainty
by: Wilson, Joey, et al.
Published: (2024)
by: Wilson, Joey, et al.
Published: (2024)
Similar Items
-
Learning Visual Feature-Based World Models via Residual Latent Action
by: Zhang, Xinyu, et al.
Published: (2026) -
Lifelong Imitation Learning with Multimodal Latent Replay and Incremental Adjustment
by: Yu, Fanqi, et al.
Published: (2026) -
UNIC: Learning Unified Multimodal Extrinsic Contact Estimation
by: Xu, Zhengtong, et al.
Published: (2026) -
EgoMimic: Scaling Imitation Learning via Egocentric Video
by: Kareer, Simar, et al.
Published: (2024) -
PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and Planning
by: Villar-Corrales, Angel, et al.
Published: (2025)