3DVLA: Enhancing Vision-Language-Action Models via 3D Spatial and Instance Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Xia, Zhongyu, Tang, Yousen, Wei, Bingqing, Wang, Yongtao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
R4Det: 4D Radar-Camera Fusion for High-Performance 3D Object Detection
by: Xia, Zhongyu, et al.
Published: (2026)
by: Xia, Zhongyu, et al.
Published: (2026)
InsFusion: Rethink Instance-level LiDAR-Camera Fusion for 3D Object Detection
by: Xia, Zhongyu, et al.
Published: (2025)
by: Xia, Zhongyu, et al.
Published: (2025)
HiDrive: A Closed-Loop Benchmark for High-Level Autonomous Driving
by: Xia, Zhongyu, et al.
Published: (2026)
by: Xia, Zhongyu, et al.
Published: (2026)
Evo-0: Vision-Language-Action Model with Implicit Spatial Understanding
by: Lin, Tao, et al.
Published: (2025)
by: Lin, Tao, et al.
Published: (2025)
Open-Set 3D Semantic Instance Maps for Vision Language Navigation -- O3D-SIM
by: Nanwani, Laksh, et al.
Published: (2024)
by: Nanwani, Laksh, et al.
Published: (2024)
SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
by: Zhang, Jiwen, et al.
Published: (2026)
by: Zhang, Jiwen, et al.
Published: (2026)
Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models
by: Yang, Yurou, et al.
Published: (2026)
by: Yang, Yurou, et al.
Published: (2026)
RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
by: Song, Chan Hee, et al.
Published: (2024)
by: Song, Chan Hee, et al.
Published: (2024)
KnowVal: A Knowledge-Augmented and Value-Guided Autonomous Driving System
by: Xia, Zhongyu, et al.
Published: (2025)
by: Xia, Zhongyu, et al.
Published: (2025)
F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions
by: Lv, Qi, et al.
Published: (2025)
by: Lv, Qi, et al.
Published: (2025)
3D CAVLA: Leveraging Depth and 3D Context to Generalize Vision Language Action Models for Unseen Tasks
by: Bhat, Vineet, et al.
Published: (2025)
by: Bhat, Vineet, et al.
Published: (2025)
VLA-R1: Enhancing Reasoning in Vision-Language-Action Models
by: Ye, Angen, et al.
Published: (2025)
by: Ye, Angen, et al.
Published: (2025)
Unified Vision-Language-Action Model
by: Wang, Yuqi, et al.
Published: (2025)
by: Wang, Yuqi, et al.
Published: (2025)
PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding
by: Song, Wenxuan, et al.
Published: (2025)
by: Song, Wenxuan, et al.
Published: (2025)
GST-VLA: Structured Gaussian Spatial Tokens for 3D Depth-Aware Vision-Language-Action Models
by: Sarowar, Md Selim, et al.
Published: (2026)
by: Sarowar, Md Selim, et al.
Published: (2026)
PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
by: Liang, Wenqi, et al.
Published: (2025)
by: Liang, Wenqi, et al.
Published: (2025)
Action Draft and Verify: A Self-Verifying Framework for Vision-Language-Action Model
by: Zhao, Chen, et al.
Published: (2026)
by: Zhao, Chen, et al.
Published: (2026)
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
by: Yang, Shuai, et al.
Published: (2025)
by: Yang, Shuai, et al.
Published: (2025)
VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
by: Sun, Jingwen, et al.
Published: (2026)
by: Sun, Jingwen, et al.
Published: (2026)
Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
by: Chen, Jiayi, et al.
Published: (2025)
by: Chen, Jiayi, et al.
Published: (2025)
Unifying 2D and 3D Vision-Language Understanding
by: Jain, Ayush, et al.
Published: (2025)
by: Jain, Ayush, et al.
Published: (2025)
SURPRISE3D: A Dataset for Spatial Understanding and Reasoning in Complex 3D Scenes
by: Huang, Jiaxin, et al.
Published: (2025)
by: Huang, Jiaxin, et al.
Published: (2025)
ZING-3D: Zero-shot Incremental 3D Scene Graphs via Vision-Language Models
by: Saxena, Pranav, et al.
Published: (2025)
by: Saxena, Pranav, et al.
Published: (2025)
VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers
by: Wang, Yating, et al.
Published: (2025)
by: Wang, Yating, et al.
Published: (2025)
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
by: Zhang, Chubin, et al.
Published: (2026)
by: Zhang, Chubin, et al.
Published: (2026)
GeneralVLA: Generalizable Vision-Language-Action Models with Knowledge-Guided Trajectory Planning
by: Ma, Guoqing, et al.
Published: (2026)
by: Ma, Guoqing, et al.
Published: (2026)
LaMP: Learning Vision-Language-Action Policies with 3D Scene Flow as Latent Motion Prior
by: Wang, Xinkai, et al.
Published: (2026)
by: Wang, Xinkai, et al.
Published: (2026)
Embodied Scene Understanding for Vision Language Models via MetaVQA
by: Wang, Weizhen, et al.
Published: (2025)
by: Wang, Weizhen, et al.
Published: (2025)
D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and Navigation
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model
by: Lin, Tao, et al.
Published: (2026)
by: Lin, Tao, et al.
Published: (2026)
PointVLA: Injecting the 3D World into Vision-Language-Action Models
by: Li, Chengmeng, et al.
Published: (2025)
by: Li, Chengmeng, et al.
Published: (2025)
FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization
by: Liu, Yicheng, et al.
Published: (2025)
by: Liu, Yicheng, et al.
Published: (2025)
Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and Tracking
by: Pätzold, Bastian, et al.
Published: (2025)
by: Pätzold, Bastian, et al.
Published: (2025)
From Monocular Vision to Autonomous Action: Guiding Tumor Resection via 3D Reconstruction
by: Acar, Ayberk, et al.
Published: (2025)
by: Acar, Ayberk, et al.
Published: (2025)
Unifying Language-Action Understanding and Generation for Autonomous Driving
by: Wang, Xinyang, et al.
Published: (2026)
by: Wang, Xinyang, et al.
Published: (2026)
FM-Fusion: Instance-aware Semantic Mapping Boosted by Vision-Language Foundation Models
by: Liu, Chuhao, et al.
Published: (2024)
by: Liu, Chuhao, et al.
Published: (2024)
Context-Nav: Context-Driven Exploration and Viewpoint-Aware 3D Spatial Reasoning for Instance Navigation
by: Jang, Won Shik, et al.
Published: (2026)
by: Jang, Won Shik, et al.
Published: (2026)
Embodied3DBench: Benchmarking Low-Level Embodied Spatial Intelligence of Vision Language Models
by: Zhang, Jiyao, et al.
Published: (2026)
by: Zhang, Jiyao, et al.
Published: (2026)
See What Matters: Differentiable Grid Sample Pruning for Generalizable Vision-Language-Action Model
by: Feng, Yixu, et al.
Published: (2026)
by: Feng, Yixu, et al.
Published: (2026)
CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
Similar Items
-
R4Det: 4D Radar-Camera Fusion for High-Performance 3D Object Detection
by: Xia, Zhongyu, et al.
Published: (2026) -
InsFusion: Rethink Instance-level LiDAR-Camera Fusion for 3D Object Detection
by: Xia, Zhongyu, et al.
Published: (2025) -
HiDrive: A Closed-Loop Benchmark for High-Level Autonomous Driving
by: Xia, Zhongyu, et al.
Published: (2026) -
Evo-0: Vision-Language-Action Model with Implicit Spatial Understanding
by: Lin, Tao, et al.
Published: (2025) -
Open-Set 3D Semantic Instance Maps for Vision Language Navigation -- O3D-SIM
by: Nanwani, Laksh, et al.
Published: (2024)