VGGT-$Ω$

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Jianyuan, Chen, Minghao, Zhang, Shangzhan, Karaev, Nikita, Schönberger, Johannes, Labatut, Patrick, Bojanowski, Piotr, Novotny, David, Vedaldi, Andrea, Rupprecht, Christian
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911685962366976
author Wang, Jianyuan
Chen, Minghao
Zhang, Shangzhan
Karaev, Nikita
Schönberger, Johannes
Labatut, Patrick
Bojanowski, Piotr
Novotny, David
Vedaldi, Andrea
Rupprecht, Christian
author_facet Wang, Jianyuan
Chen, Minghao
Zhang, Shangzhan
Karaev, Nikita
Schönberger, Johannes
Labatut, Patrick
Bojanowski, Piotr
Novotny, David
Vedaldi, Andrea
Rupprecht, Christian
contents Recent feed-forward reconstruction models, such as VGGT, have proven competitive with traditional optimization-based reconstructors while also providing geometry-aware features useful for other tasks. Here, we show that the quality of these models scales predictably with model and data size. We do so by introducing VGGT-$Ω$, which substantially improves reconstruction accuracy, efficiency, and capabilities for both static and dynamic scenes. To enable training this model at an unprecedented scale, we introduce architectural changes that improve training efficiency, a high-quality data annotation pipeline that supports dynamic scenes, and a self-supervised learning protocol. We simplify VGGT's architecture by using a single dense prediction head with multi-task supervision and removing the expensive high-resolution convolutional layers. We also use registers to aggregate scene information into a compact representation and introduce register attention, which restricts inter-frame information exchange to these registers, in part replacing global attention. In this way, during training, VGGT-$Ω$ uses only about 30% of the GPU memory of its predecessor, allowing us to train with 15x more supervised data than prior work and to leverage vast amounts of unlabeled video data. VGGT-$Ω$ achieves strong results for reconstruction of static and dynamic scenes across multiple benchmarks, for example, improving over the previous best camera estimation accuracy on Sintel by 77%. We also show that the learned registers can improve vision-language-action models and support alignment with language, suggesting that reconstruction can be a powerful and scalable proxy task for spatial understanding. Project Page: http://vggt-omega.github.io/
format Preprint
id arxiv_https___arxiv_org_abs_2605_15195
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VGGT-$Ω$
Wang, Jianyuan
Chen, Minghao
Zhang, Shangzhan
Karaev, Nikita
Schönberger, Johannes
Labatut, Patrick
Bojanowski, Piotr
Novotny, David
Vedaldi, Andrea
Rupprecht, Christian
Computer Vision and Pattern Recognition
Recent feed-forward reconstruction models, such as VGGT, have proven competitive with traditional optimization-based reconstructors while also providing geometry-aware features useful for other tasks. Here, we show that the quality of these models scales predictably with model and data size. We do so by introducing VGGT-$Ω$, which substantially improves reconstruction accuracy, efficiency, and capabilities for both static and dynamic scenes. To enable training this model at an unprecedented scale, we introduce architectural changes that improve training efficiency, a high-quality data annotation pipeline that supports dynamic scenes, and a self-supervised learning protocol. We simplify VGGT's architecture by using a single dense prediction head with multi-task supervision and removing the expensive high-resolution convolutional layers. We also use registers to aggregate scene information into a compact representation and introduce register attention, which restricts inter-frame information exchange to these registers, in part replacing global attention. In this way, during training, VGGT-$Ω$ uses only about 30% of the GPU memory of its predecessor, allowing us to train with 15x more supervised data than prior work and to leverage vast amounts of unlabeled video data. VGGT-$Ω$ achieves strong results for reconstruction of static and dynamic scenes across multiple benchmarks, for example, improving over the previous best camera estimation accuracy on Sintel by 77%. We also show that the learned registers can improve vision-language-action models and support alignment with language, suggesting that reconstruction can be a powerful and scalable proxy task for spatial understanding. Project Page: http://vggt-omega.github.io/
title VGGT-$Ω$
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.15195