DreamPlan: Efficient Reinforcement Fine-Tuning of Vision-Language Planners via Video World Models
Fuente:
arXiv
Saved in:
| Main Authors: | Jia, Emily Yue-Ting, Yuan, Weiduo, Shi, Tianheng, Guizilini, Vitor, Mao, Jiageng, Wang, Yue |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Large Reward Models: Generalizable Online Robot Reward Generation with Vision-Language Models
by: Wu, Yanru, et al.
Published: (2026)
by: Wu, Yanru, et al.
Published: (2026)
PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
by: Chow, Wei, et al.
Published: (2025)
by: Chow, Wei, et al.
Published: (2025)
Learning from Massive Human Videos for Universal Humanoid Pose Control
by: Mao, Jiageng, et al.
Published: (2024)
by: Mao, Jiageng, et al.
Published: (2024)
PlannerRFT: Reinforcing Diffusion Planners through Closed-Loop and Sample-Efficient Fine-Tuning
by: Li, Hongchen, et al.
Published: (2026)
by: Li, Hongchen, et al.
Published: (2026)
AnchorDream: Repurposing Video Diffusion for Embodiment-Aware Robot Data Synthesis
by: Ye, Junjie, et al.
Published: (2025)
by: Ye, Junjie, et al.
Published: (2025)
Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation
by: Zhao, Zhenyu, et al.
Published: (2025)
by: Zhao, Zhenyu, et al.
Published: (2025)
FuzzingRL: Reinforcement Fuzz-Testing for Revealing VLM Failures
by: Xu, Jiajun, et al.
Published: (2026)
by: Xu, Jiajun, et al.
Published: (2026)
RoboDream: Compositional World Models for Scalable Robot Data Synthesis
by: Ye, Junjie, et al.
Published: (2026)
by: Ye, Junjie, et al.
Published: (2026)
Reinforcement Learning for Adaptive Planner Parameter Tuning: A Perspective on Hierarchical Architecture
by: Wangtao, Lu, et al.
Published: (2025)
by: Wangtao, Lu, et al.
Published: (2025)
Towards Realistic Scene Generation with LiDAR Diffusion Models
by: Ran, Haoxi, et al.
Published: (2024)
by: Ran, Haoxi, et al.
Published: (2024)
Language-Augmented Symbolic Planner for Open-World Task Planning
by: Chen, Guanqi, et al.
Published: (2024)
by: Chen, Guanqi, et al.
Published: (2024)
World-aware Planning Narratives Enhance Large Vision-Language Model Planner
by: Shi, Junhao, et al.
Published: (2025)
by: Shi, Junhao, et al.
Published: (2025)
ETP-R1: Evolving Topological Planning with Reinforcement Fine-tuning for Vision-Language Navigation in Continuous Environments
by: Ye, Shuhao, et al.
Published: (2025)
by: Ye, Shuhao, et al.
Published: (2025)
Robot Learning from Any Images
by: Zhao, Siheng, et al.
Published: (2025)
by: Zhao, Siheng, et al.
Published: (2025)
Robot Learning from a Physical World Model
by: Mao, Jiageng, et al.
Published: (2025)
by: Mao, Jiageng, et al.
Published: (2025)
Fiducial Exoskeletons: Image-Centric Robot State Estimation
by: Smith, Cameron, et al.
Published: (2026)
by: Smith, Cameron, et al.
Published: (2026)
SeFA-Policy: Fast and Accurate Visuomotor Policy Learning with Selective Flow Alignment
by: Xue, Rong, et al.
Published: (2025)
by: Xue, Rong, et al.
Published: (2025)
ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation
by: Zhao, Enyu, et al.
Published: (2025)
by: Zhao, Enyu, et al.
Published: (2025)
GFM-Planner: Perception-Aware Trajectory Planning with Geometric Feature Metric
by: Lin, Yue, et al.
Published: (2025)
by: Lin, Yue, et al.
Published: (2025)
WorldRFT: Latent World Model Planning with Reinforcement Fine-Tuning for Autonomous Driving
by: Yang, Pengxuan, et al.
Published: (2025)
by: Yang, Pengxuan, et al.
Published: (2025)
ICLR: In-Context Imitation Learning with Visual Reasoning
by: Nguyen, Toan, et al.
Published: (2026)
by: Nguyen, Toan, et al.
Published: (2026)
DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
by: Zhang, Wenyao, et al.
Published: (2025)
by: Zhang, Wenyao, et al.
Published: (2025)
A Language Agent for Autonomous Driving
by: Mao, Jiageng, et al.
Published: (2023)
by: Mao, Jiageng, et al.
Published: (2023)
CO-RFT: Efficient Fine-Tuning of Vision-Language-Action Models through Chunked Offline Reinforcement Learning
by: Huang, Dongchi, et al.
Published: (2025)
by: Huang, Dongchi, et al.
Published: (2025)
Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models
by: Lyu, Mingyang, et al.
Published: (2025)
by: Lyu, Mingyang, et al.
Published: (2025)
STARE-VLA: Progressive Stage-Aware Reinforcement for Fine-Tuning Vision-Language-Action Models
by: Xu, Feng, et al.
Published: (2025)
by: Xu, Feng, et al.
Published: (2025)
$Ψ_0$: An Open Foundation Model Towards Universal Humanoid Loco-Manipulation
by: Wei, Songlin, et al.
Published: (2026)
by: Wei, Songlin, et al.
Published: (2026)
APPLV: Adaptive Planner Parameter Learning from Vision-Language-Action Model
by: Lu, Yuanjie, et al.
Published: (2026)
by: Lu, Yuanjie, et al.
Published: (2026)
BEVCALIB: LiDAR-Camera Calibration via Geometry-Guided Bird's-Eye View Representations
by: Yuan, Weiduo, et al.
Published: (2025)
by: Yuan, Weiduo, et al.
Published: (2025)
DreamToNav: Generalizable Navigation for Robots via Generative Video Planning
by: Serpiva, Valerii, et al.
Published: (2026)
by: Serpiva, Valerii, et al.
Published: (2026)
G$ \mathbf{^2} $VD Planner: Efficient Motion Planning With Grid-based Generalized Voronoi Diagrams
by: Wen, Jian, et al.
Published: (2022)
by: Wen, Jian, et al.
Published: (2022)
Driving Everywhere with Large Language Model Policy Adaptation
by: Li, Boyi, et al.
Published: (2024)
by: Li, Boyi, et al.
Published: (2024)
Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning
by: Kim, Moo Jin, et al.
Published: (2026)
by: Kim, Moo Jin, et al.
Published: (2026)
CarPlanner: Consistent Auto-regressive Trajectory Planning for Large-scale Reinforcement Learning in Autonomous Driving
by: Zhang, Dongkun, et al.
Published: (2025)
by: Zhang, Dongkun, et al.
Published: (2025)
Reinforced Embodied Planning with Verifiable Reward for Real-World Robotic Manipulation
by: Bo, Zitong, et al.
Published: (2025)
by: Bo, Zitong, et al.
Published: (2025)
MonoDream: Monocular Vision-Language Navigation with Panoramic Dreaming
by: Wang, Shuo, et al.
Published: (2025)
by: Wang, Shuo, et al.
Published: (2025)
iPlanner: Imperative Path Planning
by: Yang, Fan, et al.
Published: (2023)
by: Yang, Fan, et al.
Published: (2023)
Self-Supervised Geometry-Guided Initialization for Robust Monocular Visual Odometry
by: Kanai, Takayuki, et al.
Published: (2024)
by: Kanai, Takayuki, et al.
Published: (2024)
T3 Planner: A Self-Correcting LLM Framework for Robotic Motion Planning with Temporal Logic
by: Li, Jia, et al.
Published: (2025)
by: Li, Jia, et al.
Published: (2025)
VLM-SAFE: Vision-Language Model-Guided Safety-Aware Reinforcement Learning with World Models for Autonomous Driving
by: Qu, Yansong, et al.
Published: (2025)
by: Qu, Yansong, et al.
Published: (2025)
Similar Items
-
Large Reward Models: Generalizable Online Robot Reward Generation with Vision-Language Models
by: Wu, Yanru, et al.
Published: (2026) -
PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
by: Chow, Wei, et al.
Published: (2025) -
Learning from Massive Human Videos for Universal Humanoid Pose Control
by: Mao, Jiageng, et al.
Published: (2024) -
PlannerRFT: Reinforcing Diffusion Planners through Closed-Loop and Sample-Efficient Fine-Tuning
by: Li, Hongchen, et al.
Published: (2026) -
AnchorDream: Repurposing Video Diffusion for Embodiment-Aware Robot Data Synthesis
by: Ye, Junjie, et al.
Published: (2025)