OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Liu, Ruixun, Kong, Lingyu, Li, Derun, Zhao, Hang
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915483011252224
author Liu, Ruixun
Kong, Lingyu
Li, Derun
Zhao, Hang
author_facet Liu, Ruixun
Kong, Lingyu
Li, Derun
Zhao, Hang
contents Multimodal large language models (MLLMs) have shown strong vision-language reasoning abilities but still lack robust 3D spatial understanding, which is critical for autonomous driving. This limitation stems from two key challenges: (1) the difficulty of constructing accessible yet effective 3D representations without expensive manual annotations, and (2) the loss of fine-grained spatial details in VLMs due to the absence of large-scale 3D vision-language pretraining. To address these challenges, we propose OccVLA, a novel framework that integrates 3D occupancy representations into a unified multimodal reasoning process. Unlike prior approaches that rely on explicit 3D inputs, OccVLA treats dense 3D occupancy as both a predictive output and a supervisory signal, enabling the model to learn fine-grained spatial structures directly from 2D visual inputs. The occupancy predictions are regarded as implicit reasoning processes and can be skipped during inference without performance degradation, thereby adding no extra computational overhead. OccVLA achieves state-of-the-art results on the nuScenes benchmark for trajectory planning and demonstrates superior performance on 3D visual question-answering tasks, offering a scalable, interpretable, and fully vision-based solution for autonomous driving.
format Preprint
id arxiv_https___arxiv_org_abs_2509_05578
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision
Liu, Ruixun
Kong, Lingyu
Li, Derun
Zhao, Hang
Artificial Intelligence
Robotics
Multimodal large language models (MLLMs) have shown strong vision-language reasoning abilities but still lack robust 3D spatial understanding, which is critical for autonomous driving. This limitation stems from two key challenges: (1) the difficulty of constructing accessible yet effective 3D representations without expensive manual annotations, and (2) the loss of fine-grained spatial details in VLMs due to the absence of large-scale 3D vision-language pretraining. To address these challenges, we propose OccVLA, a novel framework that integrates 3D occupancy representations into a unified multimodal reasoning process. Unlike prior approaches that rely on explicit 3D inputs, OccVLA treats dense 3D occupancy as both a predictive output and a supervisory signal, enabling the model to learn fine-grained spatial structures directly from 2D visual inputs. The occupancy predictions are regarded as implicit reasoning processes and can be skipped during inference without performance degradation, thereby adding no extra computational overhead. OccVLA achieves state-of-the-art results on the nuScenes benchmark for trajectory planning and demonstrates superior performance on 3D visual question-answering tasks, offering a scalable, interpretable, and fully vision-based solution for autonomous driving.
title OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision
topic Artificial Intelligence
Robotics
url https://arxiv.org/abs/2509.05578