Block-Sparse Global Attention for Efficient Multi-View Geometry Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Chung-Shien Brian, Schmidt, Christian, Piekenbrinck, Jens, Leibe, Bastian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917515906514944
author Wang, Chung-Shien Brian
Schmidt, Christian
Piekenbrinck, Jens
Leibe, Bastian
author_facet Wang, Chung-Shien Brian
Schmidt, Christian
Piekenbrinck, Jens
Leibe, Bastian
contents Efficient and accurate feed-forward multi-view reconstruction has long been an important task in computer vision. Recent transformer-based models like VGGT, $π^3$ and MapAnything have demonstrated remarkable performance with relatively simple architectures. However, their scalability is fundamentally constrained by the quadratic complexity of global attention, which imposes a significant runtime bottleneck when processing large image sets. In this work, we empirically analyze the global attention matrix of these models and observe that the probability mass concentrates on a small subset of patch-patch interactions corresponding to cross-view geometric correspondences. Building on this insight and inspired by recent advances in large language models, we propose a training-free, block-sparse replacement for dense global attention, implemented with highly optimized kernels. Our method accelerates inference by more than $3\times$ while maintaining comparable task performance. Evaluations on a comprehensive suite of multi-view benchmarks demonstrate that our approach seamlessly integrates into existing global attention-based architectures such as VGGT, $π^3$ , and MapAnything, while substantially improving scalability to large image collections.
format Preprint
id arxiv_https___arxiv_org_abs_2509_07120
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Block-Sparse Global Attention for Efficient Multi-View Geometry Transformers
Wang, Chung-Shien Brian
Schmidt, Christian
Piekenbrinck, Jens
Leibe, Bastian
Computer Vision and Pattern Recognition
Efficient and accurate feed-forward multi-view reconstruction has long been an important task in computer vision. Recent transformer-based models like VGGT, $π^3$ and MapAnything have demonstrated remarkable performance with relatively simple architectures. However, their scalability is fundamentally constrained by the quadratic complexity of global attention, which imposes a significant runtime bottleneck when processing large image sets. In this work, we empirically analyze the global attention matrix of these models and observe that the probability mass concentrates on a small subset of patch-patch interactions corresponding to cross-view geometric correspondences. Building on this insight and inspired by recent advances in large language models, we propose a training-free, block-sparse replacement for dense global attention, implemented with highly optimized kernels. Our method accelerates inference by more than $3\times$ while maintaining comparable task performance. Evaluations on a comprehensive suite of multi-view benchmarks demonstrate that our approach seamlessly integrates into existing global attention-based architectures such as VGGT, $π^3$ , and MapAnything, while substantially improving scalability to large image collections.
title Block-Sparse Global Attention for Efficient Multi-View Geometry Transformers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.07120