MARVL: Multi-Stage Guidance for Robotic Manipulation via Vision-Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhou, Xunlan, Chen, Xuanlin, Zhang, Shaowei, Wan, ShengHua, Hu, Xiaohai, Yuan, Lei, Zhan, De-chuan
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914538697261056
author Zhou, Xunlan
Chen, Xuanlin
Zhang, Shaowei
Wan, ShengHua
Hu, Xiaohai
Yuan, Lei
Zhan, De-chuan
author_facet Zhou, Xunlan
Chen, Xuanlin
Zhang, Shaowei
Wan, ShengHua
Hu, Xiaohai
Yuan, Lei
Zhan, De-chuan
contents Designing dense reward functions is pivotal for efficient robotic Reinforcement Learning (RL). However, most dense rewards rely on manual engineering, which fundamentally limits the scalability and automation of reinforcement learning. While Vision-Language Models (VLMs) offer a promising path to reward design, naive VLM rewards often misalign with task progress, struggle with spatial grounding, and show limited understanding of task semantics. To address these issues, we propose MARVL-Multi-stAge guidance for Robotic manipulation via Vision-Language models. MARVL fine-tunes a VLM for spatial and semantic consistency and decomposes tasks into multi-stage subtasks with task direction projection for trajectory sensitivity. Empirically, MARVL significantly outperforms existing VLM-reward methods on the Meta-World benchmark, demonstrating superior sample efficiency and robustness on sparse-reward manipulation tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2602_15872
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MARVL: Multi-Stage Guidance for Robotic Manipulation via Vision-Language Models
Zhou, Xunlan
Chen, Xuanlin
Zhang, Shaowei
Wan, ShengHua
Hu, Xiaohai
Yuan, Lei
Zhan, De-chuan
Robotics
Computer Vision and Pattern Recognition
Machine Learning
Designing dense reward functions is pivotal for efficient robotic Reinforcement Learning (RL). However, most dense rewards rely on manual engineering, which fundamentally limits the scalability and automation of reinforcement learning. While Vision-Language Models (VLMs) offer a promising path to reward design, naive VLM rewards often misalign with task progress, struggle with spatial grounding, and show limited understanding of task semantics. To address these issues, we propose MARVL-Multi-stAge guidance for Robotic manipulation via Vision-Language models. MARVL fine-tunes a VLM for spatial and semantic consistency and decomposes tasks into multi-stage subtasks with task direction projection for trajectory sensitivity. Empirically, MARVL significantly outperforms existing VLM-reward methods on the Meta-World benchmark, demonstrating superior sample efficiency and robustness on sparse-reward manipulation tasks.
title MARVL: Multi-Stage Guidance for Robotic Manipulation via Vision-Language Models
topic Robotics
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2602.15872