SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Mengzhen, Zhou, Enshen, Chi, Cheng, Han, Yi, Rong, Shanyu, Chen, Liming, Wang, Pengwei, Wang, Zhongyuan, Zhang, Shanghang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910050705997824
author Liu, Mengzhen
Zhou, Enshen
Chi, Cheng
Han, Yi
Rong, Shanyu
Chen, Liming
Wang, Pengwei
Wang, Zhongyuan
Zhang, Shanghang
author_facet Liu, Mengzhen
Zhou, Enshen
Chi, Cheng
Han, Yi
Rong, Shanyu
Chen, Liming
Wang, Pengwei
Wang, Zhongyuan
Zhang, Shanghang
contents Active perception and manipulation are crucial for robots to interact with complex scenes. Existing methods struggle to unify semantic-driven active perception with robust, viewpoint-invariant execution. We propose SaPaVe, an end-to-end framework that jointly learns these capabilities in a data-efficient manner. Our approach decouples camera and manipulation actions rather than placing them in a shared action space, and follows a bottom-up training strategy: we first train semantic camera control on a large-scale dataset, then jointly optimize both action types using hybrid data. To support this framework, we introduce ActiveViewPose-200K, a dataset of 200k image-language-camera movement pairs for semantic camera movement learning, and a 3D geometry-aware module that improves execution robustness under dynamic viewpoints. We also present ActiveManip-Bench, the first benchmark for evaluating active manipulation beyond fixed-view settings. Extensive experiments in both simulation and real-world environments show that SaPaVe outperforms recent vision-language-action models such as GR00T N1 and \(π_0\), achieving up to 31.25\% higher success rates in real-world tasks. These results show that tightly coupled perception and execution, when trained with decoupled yet coordinated strategies, enable efficient and generalizable active manipulation. Project page: https://lmzpai.github.io/SaPaVe
format Preprint
id arxiv_https___arxiv_org_abs_2603_12193
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
Liu, Mengzhen
Zhou, Enshen
Chi, Cheng
Han, Yi
Rong, Shanyu
Chen, Liming
Wang, Pengwei
Wang, Zhongyuan
Zhang, Shanghang
Robotics
Computer Vision and Pattern Recognition
Active perception and manipulation are crucial for robots to interact with complex scenes. Existing methods struggle to unify semantic-driven active perception with robust, viewpoint-invariant execution. We propose SaPaVe, an end-to-end framework that jointly learns these capabilities in a data-efficient manner. Our approach decouples camera and manipulation actions rather than placing them in a shared action space, and follows a bottom-up training strategy: we first train semantic camera control on a large-scale dataset, then jointly optimize both action types using hybrid data. To support this framework, we introduce ActiveViewPose-200K, a dataset of 200k image-language-camera movement pairs for semantic camera movement learning, and a 3D geometry-aware module that improves execution robustness under dynamic viewpoints. We also present ActiveManip-Bench, the first benchmark for evaluating active manipulation beyond fixed-view settings. Extensive experiments in both simulation and real-world environments show that SaPaVe outperforms recent vision-language-action models such as GR00T N1 and \(π_0\), achieving up to 31.25\% higher success rates in real-world tasks. These results show that tightly coupled perception and execution, when trained with decoupled yet coordinated strategies, enable efficient and generalizable active manipulation. Project page: https://lmzpai.github.io/SaPaVe
title SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.12193