Vision-Language Model Predictive Control for Manipulation Planning and Trajectory Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Jiaming, Zhao, Wentao, Meng, Ziyu, Mao, Donghui, Song, Ran, Pan, Wei, Zhang, Wei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910905485230080
author Chen, Jiaming
Zhao, Wentao
Meng, Ziyu
Mao, Donghui
Song, Ran
Pan, Wei
Zhang, Wei
author_facet Chen, Jiaming
Zhao, Wentao
Meng, Ziyu
Mao, Donghui
Song, Ran
Pan, Wei
Zhang, Wei
contents Model Predictive Control (MPC) is a widely adopted control paradigm that leverages predictive models to estimate future system states and optimize control inputs accordingly. However, while MPC excels in planning and control, it lacks the capability for environmental perception, leading to failures in complex and unstructured scenarios. To address this limitation, we introduce Vision-Language Model Predictive Control (VLMPC), a robotic manipulation planning framework that integrates the perception power of vision-language models (VLMs) with MPC. VLMPC utilizes a conditional action sampling module that takes a goal image or language instruction as input and leverages VLM to generate candidate action sequences. These candidates are fed into a video prediction model that simulates future frames based on the actions. In addition, we propose an enhanced variant, Traj-VLMPC, which replaces video prediction with motion trajectory generation to reduce computational complexity while maintaining accuracy. Traj-VLMPC estimates motion dynamics conditioned on the candidate actions, offering a more efficient alternative for long-horizon tasks and real-time applications. Both VLMPC and Traj-VLMPC select the optimal action sequence using a VLM-based hierarchical cost function that captures both pixel-level and knowledge-level consistency between the current observation and the task input. We demonstrate that both approaches outperform existing state-of-the-art methods on public benchmarks and achieve excellent performance in various real-world robotic manipulation tasks. Code is available at https://github.com/PPjmchen/VLMPC.
format Preprint
id arxiv_https___arxiv_org_abs_2504_05225
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Vision-Language Model Predictive Control for Manipulation Planning and Trajectory Generation
Chen, Jiaming
Zhao, Wentao
Meng, Ziyu
Mao, Donghui
Song, Ran
Pan, Wei
Zhang, Wei
Robotics
Model Predictive Control (MPC) is a widely adopted control paradigm that leverages predictive models to estimate future system states and optimize control inputs accordingly. However, while MPC excels in planning and control, it lacks the capability for environmental perception, leading to failures in complex and unstructured scenarios. To address this limitation, we introduce Vision-Language Model Predictive Control (VLMPC), a robotic manipulation planning framework that integrates the perception power of vision-language models (VLMs) with MPC. VLMPC utilizes a conditional action sampling module that takes a goal image or language instruction as input and leverages VLM to generate candidate action sequences. These candidates are fed into a video prediction model that simulates future frames based on the actions. In addition, we propose an enhanced variant, Traj-VLMPC, which replaces video prediction with motion trajectory generation to reduce computational complexity while maintaining accuracy. Traj-VLMPC estimates motion dynamics conditioned on the candidate actions, offering a more efficient alternative for long-horizon tasks and real-time applications. Both VLMPC and Traj-VLMPC select the optimal action sequence using a VLM-based hierarchical cost function that captures both pixel-level and knowledge-level consistency between the current observation and the task input. We demonstrate that both approaches outperform existing state-of-the-art methods on public benchmarks and achieve excellent performance in various real-world robotic manipulation tasks. Code is available at https://github.com/PPjmchen/VLMPC.
title Vision-Language Model Predictive Control for Manipulation Planning and Trajectory Generation
topic Robotics
url https://arxiv.org/abs/2504.05225