LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token Merging

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shu, Zhijian, Lin, Cheng, Xie, Tao, Yin, Wei, Li, Ben, Pu, Zhiyuan, Li, Weize, Yao, Yao, Cao, Xun, Guo, Xiaoyang, Long, Xiao-Xiao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908694258647040
author Shu, Zhijian
Lin, Cheng
Xie, Tao
Yin, Wei
Li, Ben
Pu, Zhiyuan
Li, Weize
Yao, Yao
Cao, Xun
Guo, Xiaoyang
Long, Xiao-Xiao
author_facet Shu, Zhijian
Lin, Cheng
Xie, Tao
Yin, Wei
Li, Ben
Pu, Zhiyuan
Li, Weize
Yao, Yao
Cao, Xun
Guo, Xiaoyang
Long, Xiao-Xiao
contents 3D vision foundation models like Visual Geometry Grounded Transformer (VGGT) have advanced greatly in geometric perception. However, it is time-consuming and memory-intensive for long sequences, limiting application to large-scale scenes beyond hundreds of images. To address this, we propose LiteVGGT, achieving up to 10x speedup and substantial memory reduction, enabling efficient processing of 1000-image scenes. We derive two key insights for 3D reconstruction: (1) tokens from local image regions have inherent geometric correlations, leading to high similarity and computational redundancy; (2) token similarity across adjacent network layers remains stable, allowing for reusable merge decisions. Guided by these, we design a simple yet efficient strategy, dubbed geometry-aware cached token merging. We analyze each token's geometric importance, optimizing anchor token selection to better preserve key information for reconstruction. We also cache and reuse merge indices across layers, substantially reducing latency with minimal accuracy impact. This strategy retains VGGT's core performance, enabling efficient fine-tuning and FP8 quantization for further gains. Extensive experiments validate LiteVGGT's effectiveness, scalability, and robustness. Project page: https://garlicba.github.io/LiteVGGT/
format Preprint
id arxiv_https___arxiv_org_abs_2512_04939
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token Merging
Shu, Zhijian
Lin, Cheng
Xie, Tao
Yin, Wei
Li, Ben
Pu, Zhiyuan
Li, Weize
Yao, Yao
Cao, Xun
Guo, Xiaoyang
Long, Xiao-Xiao
Computer Vision and Pattern Recognition
3D vision foundation models like Visual Geometry Grounded Transformer (VGGT) have advanced greatly in geometric perception. However, it is time-consuming and memory-intensive for long sequences, limiting application to large-scale scenes beyond hundreds of images. To address this, we propose LiteVGGT, achieving up to 10x speedup and substantial memory reduction, enabling efficient processing of 1000-image scenes. We derive two key insights for 3D reconstruction: (1) tokens from local image regions have inherent geometric correlations, leading to high similarity and computational redundancy; (2) token similarity across adjacent network layers remains stable, allowing for reusable merge decisions. Guided by these, we design a simple yet efficient strategy, dubbed geometry-aware cached token merging. We analyze each token's geometric importance, optimizing anchor token selection to better preserve key information for reconstruction. We also cache and reuse merge indices across layers, substantially reducing latency with minimal accuracy impact. This strategy retains VGGT's core performance, enabling efficient fine-tuning and FP8 quantization for further gains. Extensive experiments validate LiteVGGT's effectiveness, scalability, and robustness. Project page: https://garlicba.github.io/LiteVGGT/
title LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token Merging
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.04939