OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Peng, Haosong, Li, Hao, Dai, Yalun, Lan, Yushi, Luo, Yihang, Qi, Tianyu, Zhang, Zhengshen, Zhan, Yufeng, Zhang, Junfei, Xu, Wenchao, Liu, Ziwei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911265731903488
author Peng, Haosong
Li, Hao
Dai, Yalun
Lan, Yushi
Luo, Yihang
Qi, Tianyu
Zhang, Zhengshen
Zhan, Yufeng
Zhang, Junfei
Xu, Wenchao
Liu, Ziwei
author_facet Peng, Haosong
Li, Hao
Dai, Yalun
Lan, Yushi
Luo, Yihang
Qi, Tianyu
Zhang, Zhengshen
Zhan, Yufeng
Zhang, Junfei
Xu, Wenchao
Liu, Ziwei
contents General 3D foundation models have started to lead the trend of unifying diverse vision tasks, yet most assume RGB-only inputs and ignore readily available geometric cues (e.g., camera intrinsics, poses, and depth maps). To address this issue, we introduce OmniVGGT, a novel framework that can effectively benefit from an arbitrary number of auxiliary geometric modalities during both training and inference. In our framework, a GeoAdapter is proposed to encode depth and camera intrinsics/extrinsics into a spatial foundation model. It employs zero-initialized convolutions to progressively inject geometric information without disrupting the foundation model's representation space. This design ensures stable optimization with negligible overhead, maintaining inference speed comparable to VGGT even with multiple additional inputs. Additionally, a stochastic multimodal fusion regimen is proposed, which randomly samples modality subsets per instance during training. This enables an arbitrary number of modality inputs during testing and promotes learning robust spatial representations instead of overfitting to auxiliary cues. Comprehensive experiments on monocular/multi-view depth estimation, multi-view stereo, and camera pose estimation demonstrate that OmniVGGT outperforms prior methods with auxiliary inputs and achieves state-of-the-art results even with RGB-only input. To further highlight its practical utility, we integrated OmniVGGT into vision-language-action (VLA) models. The enhanced VLA model by OmniVGGT not only outperforms the vanilla point-cloud-based baseline on mainstream benchmarks, but also effectively leverages accessible auxiliary inputs to achieve consistent gains on robotic tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2511_10560
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer
Peng, Haosong
Li, Hao
Dai, Yalun
Lan, Yushi
Luo, Yihang
Qi, Tianyu
Zhang, Zhengshen
Zhan, Yufeng
Zhang, Junfei
Xu, Wenchao
Liu, Ziwei
Computer Vision and Pattern Recognition
General 3D foundation models have started to lead the trend of unifying diverse vision tasks, yet most assume RGB-only inputs and ignore readily available geometric cues (e.g., camera intrinsics, poses, and depth maps). To address this issue, we introduce OmniVGGT, a novel framework that can effectively benefit from an arbitrary number of auxiliary geometric modalities during both training and inference. In our framework, a GeoAdapter is proposed to encode depth and camera intrinsics/extrinsics into a spatial foundation model. It employs zero-initialized convolutions to progressively inject geometric information without disrupting the foundation model's representation space. This design ensures stable optimization with negligible overhead, maintaining inference speed comparable to VGGT even with multiple additional inputs. Additionally, a stochastic multimodal fusion regimen is proposed, which randomly samples modality subsets per instance during training. This enables an arbitrary number of modality inputs during testing and promotes learning robust spatial representations instead of overfitting to auxiliary cues. Comprehensive experiments on monocular/multi-view depth estimation, multi-view stereo, and camera pose estimation demonstrate that OmniVGGT outperforms prior methods with auxiliary inputs and achieves state-of-the-art results even with RGB-only input. To further highlight its practical utility, we integrated OmniVGGT into vision-language-action (VLA) models. The enhanced VLA model by OmniVGGT not only outperforms the vanilla point-cloud-based baseline on mainstream benchmarks, but also effectively leverages accessible auxiliary inputs to achieve consistent gains on robotic tasks.
title OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.10560