Spark3R: Asymmetric Token Reduction Makes Fast Feed-Forward 3D Reconstruction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Zecheng, Fu, Jiaye, Gao, Qiankun, Li, Haijie, Wu, Yanmin, Zhang, Jiaqi, Ma, Siwei, Zhang, Jian
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918509983825920
author Tang, Zecheng
Fu, Jiaye
Gao, Qiankun
Li, Haijie
Wu, Yanmin
Zhang, Jiaqi
Ma, Siwei
Zhang, Jian
author_facet Tang, Zecheng
Fu, Jiaye
Gao, Qiankun
Li, Haijie
Wu, Yanmin
Zhang, Jiaqi
Ma, Siwei
Zhang, Jian
contents Feed-forward 3D reconstruction models based on Vision Transformers can directly estimate scene geometry and camera poses from a small set of input images, but scaling them to video inputs with hundreds or thousands of frames remains challenging due to the quadratic cost of global attention layers. Recent token-merging methods accelerate these models by compressing the token sequence within the global attention layers, but they apply a uniform reduction to query tokens and key-value tokens, ignoring their functionally distinct roles in 3D reconstruction. In this work, we identify a key property of feed-forward 3D reconstruction models: query tokens encode view-specific geometric requests and are sensitive to compression, while key-value tokens represent shared scene context and tolerate aggressive compression. Guided by this insight, we propose Spark3R, a training-free acceleration framework that decouples the compression of query tokens and key-value tokens by assigning distinct reduction factors, with intra-group token merging applied to query tokens and lightweight token pruning to key-value tokens. Additionally, Spark3R adaptively adjusts the key-value reduction factor across layers, further improving the quality-efficiency trade-off. As a plug-and-play framework requiring no retraining, Spark3R integrates directly into multiple pretrained feed-forward 3D reconstruction models, including VGGT, $π^3$, Depth-Anything-3, and VGGT-$Ω$, and achieves up to $28\times$ speedup on 1,000-frame inputs while maintaining competitive reconstruction quality.
format Preprint
id arxiv_https___arxiv_org_abs_2605_06270
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Spark3R: Asymmetric Token Reduction Makes Fast Feed-Forward 3D Reconstruction
Tang, Zecheng
Fu, Jiaye
Gao, Qiankun
Li, Haijie
Wu, Yanmin
Zhang, Jiaqi
Ma, Siwei
Zhang, Jian
Computer Vision and Pattern Recognition
Feed-forward 3D reconstruction models based on Vision Transformers can directly estimate scene geometry and camera poses from a small set of input images, but scaling them to video inputs with hundreds or thousands of frames remains challenging due to the quadratic cost of global attention layers. Recent token-merging methods accelerate these models by compressing the token sequence within the global attention layers, but they apply a uniform reduction to query tokens and key-value tokens, ignoring their functionally distinct roles in 3D reconstruction. In this work, we identify a key property of feed-forward 3D reconstruction models: query tokens encode view-specific geometric requests and are sensitive to compression, while key-value tokens represent shared scene context and tolerate aggressive compression. Guided by this insight, we propose Spark3R, a training-free acceleration framework that decouples the compression of query tokens and key-value tokens by assigning distinct reduction factors, with intra-group token merging applied to query tokens and lightweight token pruning to key-value tokens. Additionally, Spark3R adaptively adjusts the key-value reduction factor across layers, further improving the quality-efficiency trade-off. As a plug-and-play framework requiring no retraining, Spark3R integrates directly into multiple pretrained feed-forward 3D reconstruction models, including VGGT, $π^3$, Depth-Anything-3, and VGGT-$Ω$, and achieves up to $28\times$ speedup on 1,000-frame inputs while maintaining competitive reconstruction quality.
title Spark3R: Asymmetric Token Reduction Makes Fast Feed-Forward 3D Reconstruction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.06270