DVGT-2: Vision-Geometry-Action Model for Autonomous Driving at Scale

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zuo, Sicheng, Xie, Zixun, Zheng, Wenzhao, Xu, Shaoqing, Li, Fang, Li, Hanbing, Chen, Long, Yang, Zhi-Xin, Lu, Jiwen
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913058855583744
author Zuo, Sicheng
Xie, Zixun
Zheng, Wenzhao
Xu, Shaoqing
Li, Fang
Li, Hanbing
Chen, Long
Yang, Zhi-Xin
Lu, Jiwen
author_facet Zuo, Sicheng
Xie, Zixun
Zheng, Wenzhao
Xu, Shaoqing
Li, Fang
Li, Hanbing
Chen, Long
Yang, Zhi-Xin
Lu, Jiwen
contents End-to-end autonomous driving has evolved from the conventional paradigm based on sparse perception into vision-language-action (VLA) models, which focus on learning language descriptions as an auxiliary task to facilitate planning. In this paper, we propose an alternative Vision-Geometry-Action (VGA) paradigm that advocates dense 3D geometry as the critical cue for autonomous driving. As vehicles operate in a 3D world, we think dense 3D geometry provides the most comprehensive information for decision-making. However, most existing geometry reconstruction methods (e.g., DVGT) rely on computationally expensive batch processing of multi-frame inputs and cannot be applied to online planning. To address this, we introduce a streaming Driving Visual Geometry Transformer (DVGT-2), which processes inputs in an online manner and jointly outputs dense geometry and trajectory planning for the current frame. We employ temporal causal attention and cache historical features to support on-the-fly inference. To further enhance efficiency, we propose a sliding-window streaming strategy and use historical caches within a certain interval to avoid repetitive computations. Despite the faster speed, DVGT-2 achieves superior geometry reconstruction performance on various datasets. The same trained DVGT-2 can be directly applied to planning across diverse camera configurations without fine-tuning, including closed-loop NAVSIM and open-loop nuScenes benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2604_00813
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DVGT-2: Vision-Geometry-Action Model for Autonomous Driving at Scale
Zuo, Sicheng
Xie, Zixun
Zheng, Wenzhao
Xu, Shaoqing
Li, Fang
Li, Hanbing
Chen, Long
Yang, Zhi-Xin
Lu, Jiwen
Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
End-to-end autonomous driving has evolved from the conventional paradigm based on sparse perception into vision-language-action (VLA) models, which focus on learning language descriptions as an auxiliary task to facilitate planning. In this paper, we propose an alternative Vision-Geometry-Action (VGA) paradigm that advocates dense 3D geometry as the critical cue for autonomous driving. As vehicles operate in a 3D world, we think dense 3D geometry provides the most comprehensive information for decision-making. However, most existing geometry reconstruction methods (e.g., DVGT) rely on computationally expensive batch processing of multi-frame inputs and cannot be applied to online planning. To address this, we introduce a streaming Driving Visual Geometry Transformer (DVGT-2), which processes inputs in an online manner and jointly outputs dense geometry and trajectory planning for the current frame. We employ temporal causal attention and cache historical features to support on-the-fly inference. To further enhance efficiency, we propose a sliding-window streaming strategy and use historical caches within a certain interval to avoid repetitive computations. Despite the faster speed, DVGT-2 achieves superior geometry reconstruction performance on various datasets. The same trained DVGT-2 can be directly applied to planning across diverse camera configurations without fine-tuning, including closed-loop NAVSIM and open-loop nuScenes benchmarks.
title DVGT-2: Vision-Geometry-Action Model for Autonomous Driving at Scale
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2604.00813