StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Deng, Shengliang, Yan, Mi, Zheng, Yixin, Su, Jiayi, Zhang, Wenhao, Zhao, Xiaoguang, Cui, Heming, Zhang, Zhizheng, Wang, He
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915694766981120
author Deng, Shengliang
Yan, Mi
Zheng, Yixin
Su, Jiayi
Zhang, Wenhao
Zhao, Xiaoguang
Cui, Heming
Zhang, Zhizheng
Wang, He
author_facet Deng, Shengliang
Yan, Mi
Zheng, Yixin
Su, Jiayi
Zhang, Wenhao
Zhao, Xiaoguang
Cui, Heming
Zhang, Zhizheng
Wang, He
contents Stereo cameras closely mimic human binocular vision, providing rich spatial cues critical for precise robotic manipulation. Despite their advantage, the adoption of stereo vision in vision-language-action models (VLAs) remains underexplored. In this work, we present StereoVLA, a VLA model that leverages rich geometric cues from stereo vision. We propose a novel Geometric-Semantic Feature Extraction module that utilizes vision foundation models to extract and fuse two key features: 1) geometric features from subtle stereo-view differences for spatial perception; 2) semantic-rich features from the monocular view for instruction following. Additionally, we propose an auxiliary Interaction-Region Depth Estimation task to further enhance spatial perception and accelerate model convergence. Extensive experiments show that our approach outperforms baselines by a large margin in diverse tasks under the stereo setting and demonstrates strong robustness to camera pose variations.
format Preprint
id arxiv_https___arxiv_org_abs_2512_21970
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
Deng, Shengliang
Yan, Mi
Zheng, Yixin
Su, Jiayi
Zhang, Wenhao
Zhao, Xiaoguang
Cui, Heming
Zhang, Zhizheng
Wang, He
Robotics
Stereo cameras closely mimic human binocular vision, providing rich spatial cues critical for precise robotic manipulation. Despite their advantage, the adoption of stereo vision in vision-language-action models (VLAs) remains underexplored. In this work, we present StereoVLA, a VLA model that leverages rich geometric cues from stereo vision. We propose a novel Geometric-Semantic Feature Extraction module that utilizes vision foundation models to extract and fuse two key features: 1) geometric features from subtle stereo-view differences for spatial perception; 2) semantic-rich features from the monocular view for instruction following. Additionally, we propose an auxiliary Interaction-Region Depth Estimation task to further enhance spatial perception and accelerate model convergence. Extensive experiments show that our approach outperforms baselines by a large margin in diverse tasks under the stereo setting and demonstrates strong robustness to camera pose variations.
title StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
topic Robotics
url https://arxiv.org/abs/2512.21970