PaceVGGT: Pre-Alternating-Attention Token Pruning for Visual Geometry Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Haotang, Qi, Zhenyu, Wang, Shaohan Henry, Peng, Kebin, Wang, Zi, Guo, Qing, He, Sen, Yang, Huanrui
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909028317134848
author Li, Haotang
Qi, Zhenyu
Wang, Shaohan Henry
Peng, Kebin
Wang, Zi
Guo, Qing
He, Sen
Yang, Huanrui
author_facet Li, Haotang
Qi, Zhenyu
Wang, Shaohan Henry
Peng, Kebin
Wang, Zi
Guo, Qing
He, Sen
Yang, Huanrui
contents Visual Geometry Transformer (VGGT) is a strong feed-forward model for multiple 3D tasks, but its Alternating-Attention (AA) stack scales quadratically in the total token count, making long clips expensive. Existing token-reduction accelerators operate inside AA, leaving the patch grid that enters AA uncompressed. We introduce PaceVGGT, a pre-AA token pruning framework that prunes DINO patch tokens before the first AA block of a frozen VGGT. PaceVGGT trains a lightweight Token Scorer that estimates per-token importance from DINO features. The scorer is first distilled against an AA-internal attention target from the unpruned backbone, then refined under downstream camera, depth, and point-map losses. A per-frame keep budget fixes the backbone-visible sequence length, while an importance-adaptive merge/prune assignment preserves residual content from high-saliency frames under a fixed total merge budget. A Feature-guided Restoration module reconstructs the dense spatial grid required by the prediction heads. On ScanNet-50 and 7-Scenes, PaceVGGT remains on the reconstruction quality--latency frontier while reducing inference latency. On ScanNet-50, it reduces latency by \(5.1\times\) over unmodified VGGT at \(N=300\) and \(1.47\times\) over LiteVGGT at \(N=1000\). These results identify pre-AA pruning as a viable acceleration route for frozen VGGT-style geometry transformers.
format Preprint
id arxiv_https___arxiv_org_abs_2605_08371
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PaceVGGT: Pre-Alternating-Attention Token Pruning for Visual Geometry Transformers
Li, Haotang
Qi, Zhenyu
Wang, Shaohan Henry
Peng, Kebin
Wang, Zi
Guo, Qing
He, Sen
Yang, Huanrui
Computer Vision and Pattern Recognition
Visual Geometry Transformer (VGGT) is a strong feed-forward model for multiple 3D tasks, but its Alternating-Attention (AA) stack scales quadratically in the total token count, making long clips expensive. Existing token-reduction accelerators operate inside AA, leaving the patch grid that enters AA uncompressed. We introduce PaceVGGT, a pre-AA token pruning framework that prunes DINO patch tokens before the first AA block of a frozen VGGT. PaceVGGT trains a lightweight Token Scorer that estimates per-token importance from DINO features. The scorer is first distilled against an AA-internal attention target from the unpruned backbone, then refined under downstream camera, depth, and point-map losses. A per-frame keep budget fixes the backbone-visible sequence length, while an importance-adaptive merge/prune assignment preserves residual content from high-saliency frames under a fixed total merge budget. A Feature-guided Restoration module reconstructs the dense spatial grid required by the prediction heads. On ScanNet-50 and 7-Scenes, PaceVGGT remains on the reconstruction quality--latency frontier while reducing inference latency. On ScanNet-50, it reduces latency by \(5.1\times\) over unmodified VGGT at \(N=300\) and \(1.47\times\) over LiteVGGT at \(N=1000\). These results identify pre-AA pruning as a viable acceleration route for frozen VGGT-style geometry transformers.
title PaceVGGT: Pre-Alternating-Attention Token Pruning for Visual Geometry Transformers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.08371