Streaming 4D Visual Geometry Transformer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhuo, Dong, Zheng, Wenzhao, Guo, Jiahe, Wu, Yuqi, Zhou, Jie, Lu, Jiwen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915900407414784
author Zhuo, Dong
Zheng, Wenzhao
Guo, Jiahe
Wu, Yuqi
Zhou, Jie
Lu, Jiwen
author_facet Zhuo, Dong
Zheng, Wenzhao
Guo, Jiahe
Wu, Yuqi
Zhou, Jie
Lu, Jiwen
contents Perceiving and reconstructing 3D geometry from videos is a fundamental yet challenging computer vision task. To facilitate interactive and low-latency applications, we propose a streaming visual geometry transformer that shares a similar philosophy with autoregressive large language models. We explore a simple and efficient design and employ a causal transformer architecture to process the input sequence in an online manner. We use temporal causal attention and cache the historical keys and values as implicit memory to enable efficient streaming long-term 3D reconstruction. This design can handle low-latency 3D reconstruction by incrementally integrating historical information while maintaining high-quality spatial consistency. For efficient training, we propose to distill knowledge from the dense bidirectional visual geometry grounded transformer (VGGT) to our causal model. For inference, our model supports the migration of optimized efficient attention operators (e.g., FlashAttention) from large language models. Extensive experiments on various 3D geometry perception benchmarks demonstrate that our model enhances inference speed in online scenarios while maintaining competitive performance, thereby facilitating scalable and interactive 3D vision systems. Code is available at: https://github.com/wzzheng/StreamVGGT.
format Preprint
id arxiv_https___arxiv_org_abs_2507_11539
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Streaming 4D Visual Geometry Transformer
Zhuo, Dong
Zheng, Wenzhao
Guo, Jiahe
Wu, Yuqi
Zhou, Jie
Lu, Jiwen
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Perceiving and reconstructing 3D geometry from videos is a fundamental yet challenging computer vision task. To facilitate interactive and low-latency applications, we propose a streaming visual geometry transformer that shares a similar philosophy with autoregressive large language models. We explore a simple and efficient design and employ a causal transformer architecture to process the input sequence in an online manner. We use temporal causal attention and cache the historical keys and values as implicit memory to enable efficient streaming long-term 3D reconstruction. This design can handle low-latency 3D reconstruction by incrementally integrating historical information while maintaining high-quality spatial consistency. For efficient training, we propose to distill knowledge from the dense bidirectional visual geometry grounded transformer (VGGT) to our causal model. For inference, our model supports the migration of optimized efficient attention operators (e.g., FlashAttention) from large language models. Extensive experiments on various 3D geometry perception benchmarks demonstrate that our model enhances inference speed in online scenarios while maintaining competitive performance, thereby facilitating scalable and interactive 3D vision systems. Code is available at: https://github.com/wzzheng/StreamVGGT.
title Streaming 4D Visual Geometry Transformer
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2507.11539