DVGT: Driving Visual Geometry Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | Zuo, Sicheng, Xie, Zixun, Zheng, Wenzhao, Xu, Shaoqing, Li, Fang, Jiang, Shengyin, Chen, Long, Yang, Zhi-Xin, Lu, Jiwen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DVGT-2: Vision-Geometry-Action Model for Autonomous Driving at Scale
by: Zuo, Sicheng, et al.
Published: (2026)
by: Zuo, Sicheng, et al.
Published: (2026)
GPD-1: Generative Pre-training for Driving
by: Xie, Zixun, et al.
Published: (2024)
by: Xie, Zixun, et al.
Published: (2024)
Vega: Learning to Drive with Natural Language Instructions
by: Zuo, Sicheng, et al.
Published: (2026)
by: Zuo, Sicheng, et al.
Published: (2026)
GaussianAD: Gaussian-Centric End-to-End Autonomous Driving
by: Zheng, Wenzhao, et al.
Published: (2024)
by: Zheng, Wenzhao, et al.
Published: (2024)
GaussianPretrain: A Simple Unified 3D Gaussian Representation for Visual Pre-training in Autonomous Driving
by: Xu, Shaoqing, et al.
Published: (2024)
by: Xu, Shaoqing, et al.
Published: (2024)
VILTA: A VLM-in-the-Loop Adversary for Enhancing Driving Policy Robustness
by: Chen, Qimao, et al.
Published: (2026)
by: Chen, Qimao, et al.
Published: (2026)
AdaThinkDrive: Adaptive Thinking via Reinforcement Learning for Autonomous Driving
by: Luo, Yuechen, et al.
Published: (2025)
by: Luo, Yuechen, et al.
Published: (2025)
DriveTok: 3D Driving Scene Tokenization for Unified Multi-View Reconstruction and Understanding
by: Zhuo, Dong, et al.
Published: (2026)
by: Zhuo, Dong, et al.
Published: (2026)
GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control
by: Chen, Anthony, et al.
Published: (2025)
by: Chen, Anthony, et al.
Published: (2025)
Doe-1: Closed-Loop Autonomous Driving with Large World Model
by: Zheng, Wenzhao, et al.
Published: (2024)
by: Zheng, Wenzhao, et al.
Published: (2024)
Streaming 4D Visual Geometry Transformer
by: Zhuo, Dong, et al.
Published: (2025)
by: Zhuo, Dong, et al.
Published: (2025)
QuadricFormer: Scene as Superquadrics for 3D Semantic Occupancy Prediction
by: Zuo, Sicheng, et al.
Published: (2025)
by: Zuo, Sicheng, et al.
Published: (2025)
GaussianWorld: Gaussian World Model for Streaming 3D Occupancy Prediction
by: Zuo, Sicheng, et al.
Published: (2024)
by: Zuo, Sicheng, et al.
Published: (2024)
AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation
by: Guo, Wenxuan, et al.
Published: (2026)
by: Guo, Wenxuan, et al.
Published: (2026)
EmbodiedOcc: Embodied 3D Occupancy Prediction for Vision-based Online Scene Understanding
by: Wu, Yuqi, et al.
Published: (2024)
by: Wu, Yuqi, et al.
Published: (2024)
DIVER: Reinforced Diffusion Breaks Imitation Bottlenecks in End-to-End Autonomous Driving
by: Song, Ziying, et al.
Published: (2025)
by: Song, Ziying, et al.
Published: (2025)
HawkDrive: A Transformer-driven Visual Perception System for Autonomous Driving in Night Scene
by: Guo, Ziang, et al.
Published: (2024)
by: Guo, Ziang, et al.
Published: (2024)
Look, Zoom, Understand: The Robotic Eyeball for Embodied Perception
by: Yang, Jiashu, et al.
Published: (2025)
by: Yang, Jiashu, et al.
Published: (2025)
LIDEA: Human-to-Robot Imitation Learning via Implicit Feature Distillation and Explicit Geometry Alignment
by: Xu, Yifu, et al.
Published: (2026)
by: Xu, Yifu, et al.
Published: (2026)
SG-Nav: Online 3D Scene Graph Prompting for LLM-based Zero-shot Object Navigation
by: Yin, Hang, et al.
Published: (2024)
by: Yin, Hang, et al.
Published: (2024)
R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation
by: Xu, Xiuwei, et al.
Published: (2025)
by: Xu, Xiuwei, et al.
Published: (2025)
DriveCode: Domain Specific Numerical Encoding for LLM-Based Autonomous Driving
by: Wang, Zhiye, et al.
Published: (2026)
by: Wang, Zhiye, et al.
Published: (2026)
GDRNPP: A Geometry-guided and Fully Learning-based Object Pose Estimator
by: Liu, Xingyu, et al.
Published: (2021)
by: Liu, Xingyu, et al.
Published: (2021)
U-ViLAR: Uncertainty-Aware Visual Localization for Autonomous Driving via Differentiable Association and Registration
by: Li, Xiaofan, et al.
Published: (2025)
by: Li, Xiaofan, et al.
Published: (2025)
GMF-Drive: Gated Mamba Fusion with Spatial-Aware BEV Representation for End-to-End Autonomous Driving
by: Wang, Jian, et al.
Published: (2025)
by: Wang, Jian, et al.
Published: (2025)
The RoboDrive Challenge: Drive Anytime Anywhere in Any Condition
by: Kong, Lingdong, et al.
Published: (2024)
by: Kong, Lingdong, et al.
Published: (2024)
UniGoal: Towards Universal Zero-shot Goal-oriented Navigation
by: Yin, Hang, et al.
Published: (2025)
by: Yin, Hang, et al.
Published: (2025)
EmbodiedSAM: Online Segment Any 3D Thing in Real Time
by: Xu, Xiuwei, et al.
Published: (2024)
by: Xu, Xiuwei, et al.
Published: (2024)
GC-VLN: Instruction as Graph Constraints for Training-free Vision-and-Language Navigation
by: Yin, Hang, et al.
Published: (2025)
by: Yin, Hang, et al.
Published: (2025)
V2M: Visual 2-Dimensional Mamba for Image Representation Learning
by: Wang, Chengkun, et al.
Published: (2024)
by: Wang, Chengkun, et al.
Published: (2024)
A Model-based Visual Contact Localization and Force Sensing System for Compliant Robotic Grippers
by: Zuo, Kaiwen, et al.
Published: (2026)
by: Zuo, Kaiwen, et al.
Published: (2026)
Joint 3D Geometry Reconstruction and Motion Generation for 4D Synthesis from a Single Image
by: Zhang, Yanran, et al.
Published: (2025)
by: Zhang, Yanran, et al.
Published: (2025)
OmniIndoor3D: Comprehensive Indoor 3D Reconstruction
by: Wei, Xiaobao, et al.
Published: (2025)
by: Wei, Xiaobao, et al.
Published: (2025)
Driving Intents Amplify Planning-Oriented Reinforcement Learning
by: Lu, Hengtong, et al.
Published: (2026)
by: Lu, Hengtong, et al.
Published: (2026)
ManiGaussian: Dynamic Gaussian Splatting for Multi-task Robotic Manipulation
by: Lu, Guanxing, et al.
Published: (2024)
by: Lu, Guanxing, et al.
Published: (2024)
LSGS-Loc: Towards Robust 3DGS-Based Visual Localization for Large-Scale UAV Scenarios
by: Zhang, Xiang, et al.
Published: (2026)
by: Zhang, Xiang, et al.
Published: (2026)
ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving
by: Li, Yongkang, et al.
Published: (2025)
by: Li, Yongkang, et al.
Published: (2025)
Is Your Driving World Model an All-Around Player?
by: Kong, Lingdong, et al.
Published: (2026)
by: Kong, Lingdong, et al.
Published: (2026)
TiGDistill-BEV: Multi-view BEV 3D Object Detection via Target Inner-Geometry Learning Distillation
by: Xu, Shaoqing, et al.
Published: (2024)
by: Xu, Shaoqing, et al.
Published: (2024)
GeoPredict: Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA Manipulation
by: Qian, Jingjing, et al.
Published: (2025)
by: Qian, Jingjing, et al.
Published: (2025)
Similar Items
-
DVGT-2: Vision-Geometry-Action Model for Autonomous Driving at Scale
by: Zuo, Sicheng, et al.
Published: (2026) -
GPD-1: Generative Pre-training for Driving
by: Xie, Zixun, et al.
Published: (2024) -
Vega: Learning to Drive with Natural Language Instructions
by: Zuo, Sicheng, et al.
Published: (2026) -
GaussianAD: Gaussian-Centric End-to-End Autonomous Driving
by: Zheng, Wenzhao, et al.
Published: (2024) -
GaussianPretrain: A Simple Unified 3D Gaussian Representation for Visual Pre-training in Autonomous Driving
by: Xu, Shaoqing, et al.
Published: (2024)