Analyzing the Mechanism of Attention Collapse in VGGT from a Dynamics Perspective
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Huan, Luo, Longjun, Shi, Yuling, Gu, Xiaodong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VGGT-X: When VGGT Meets Dense Novel View Synthesis
von: Liu, Yang, et al.
Veröffentlicht: (2025)
von: Liu, Yang, et al.
Veröffentlicht: (2025)
AVGGT: Rethinking Global Attention for Accelerating VGGT
von: Sun, Xianbing, et al.
Veröffentlicht: (2025)
von: Sun, Xianbing, et al.
Veröffentlicht: (2025)
VGGT-World: Transforming VGGT into an Autoregressive Geometry World Model
von: Sun, Xiangyu, et al.
Veröffentlicht: (2026)
von: Sun, Xiangyu, et al.
Veröffentlicht: (2026)
VGGT-$Ω$
von: Wang, Jianyuan, et al.
Veröffentlicht: (2026)
von: Wang, Jianyuan, et al.
Veröffentlicht: (2026)
PaceVGGT: Pre-Alternating-Attention Token Pruning for Visual Geometry Transformers
von: Li, Haotang, et al.
Veröffentlicht: (2026)
von: Li, Haotang, et al.
Veröffentlicht: (2026)
LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token Merging
von: Shu, Zhijian, et al.
Veröffentlicht: (2025)
von: Shu, Zhijian, et al.
Veröffentlicht: (2025)
TurboVGGT: Fast Visual Geometry Reconstruction with Adaptive Alternating Attention
von: Huang, David, et al.
Veröffentlicht: (2026)
von: Huang, David, et al.
Veröffentlicht: (2026)
PanoVGGT: Feed-Forward 3D Reconstruction from Panoramic Imagery
von: Guo, Yijing, et al.
Veröffentlicht: (2026)
von: Guo, Yijing, et al.
Veröffentlicht: (2026)
VGGT-MPR: VGGT-Enhanced Multimodal Place Recognition in Autonomous Driving Environments
von: Xu, Jingyi, et al.
Veröffentlicht: (2026)
von: Xu, Jingyi, et al.
Veröffentlicht: (2026)
FrameVGGT: Geometry-Aligned Frame-Level Memory for Bounded Streaming VGGT
von: Xu, Zhisong, et al.
Veröffentlicht: (2026)
von: Xu, Zhisong, et al.
Veröffentlicht: (2026)
FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor Attention
von: Wang, Zipeng, et al.
Veröffentlicht: (2025)
von: Wang, Zipeng, et al.
Veröffentlicht: (2025)
VGGT-SLAM++
von: Mandal, Avilasha, et al.
Veröffentlicht: (2026)
von: Mandal, Avilasha, et al.
Veröffentlicht: (2026)
DynamicVGGT: Learning Dynamic Point Maps for 4D Scene Reconstruction in Autonomous Driving
von: He, Zhuolin, et al.
Veröffentlicht: (2026)
von: He, Zhuolin, et al.
Veröffentlicht: (2026)
VGGT-Long: Chunk it, Loop it, Align it -- Pushing VGGT's Limits on Kilometer-scale Long RGB Sequences
von: Deng, Kai, et al.
Veröffentlicht: (2025)
von: Deng, Kai, et al.
Veröffentlicht: (2025)
VGGT: Visual Geometry Grounded Transformer
von: Wang, Jianyuan, et al.
Veröffentlicht: (2025)
von: Wang, Jianyuan, et al.
Veröffentlicht: (2025)
Dense Semantic Matching with VGGT Prior
von: Yang, Songlin, et al.
Veröffentlicht: (2025)
von: Yang, Songlin, et al.
Veröffentlicht: (2025)
OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer
von: Peng, Haosong, et al.
Veröffentlicht: (2025)
von: Peng, Haosong, et al.
Veröffentlicht: (2025)
Progressive Supernet Training for Efficient Visual Autoregressive Modeling
von: Chen, Xiaoyue, et al.
Veröffentlicht: (2025)
von: Chen, Xiaoyue, et al.
Veröffentlicht: (2025)
HD-VGGT: High-Resolution Visual Geometry Transformer
von: Chen, Tianrun, et al.
Veröffentlicht: (2026)
von: Chen, Tianrun, et al.
Veröffentlicht: (2026)
VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection
von: Cao, Yang, et al.
Veröffentlicht: (2026)
von: Cao, Yang, et al.
Veröffentlicht: (2026)
EventVGGT: Exploring Cross-Modal Distillation for Consistent Event-based Depth Estimation
von: Ren, Yinrui, et al.
Veröffentlicht: (2026)
von: Ren, Yinrui, et al.
Veröffentlicht: (2026)
GPA-VGGT:Adapting VGGT to Large Scale Localization by Self-Supervised Learning with Geometry and Physics Aware Loss
von: Xu, Yangfan, et al.
Veröffentlicht: (2026)
von: Xu, Yangfan, et al.
Veröffentlicht: (2026)
Sparsity-Aware Voxel Attention and Foreground Modulation for 3D Semantic Scene Completion
von: Xue, Yu, et al.
Veröffentlicht: (2026)
von: Xue, Yu, et al.
Veröffentlicht: (2026)
VGGT-Segmentor: Geometry-Enhanced Cross-View Segmentation
von: Gao, Yulu, et al.
Veröffentlicht: (2026)
von: Gao, Yulu, et al.
Veröffentlicht: (2026)
SegVGGT: Joint 3D Reconstruction and Instance Segmentation from Multi-View Images
von: Qu, Jinyuan, et al.
Veröffentlicht: (2026)
von: Qu, Jinyuan, et al.
Veröffentlicht: (2026)
S-VGGT: Structure-Aware Subscene Decomposition for Scalable 3D Foundation Models
von: Li, Xinze, et al.
Veröffentlicht: (2026)
von: Li, Xinze, et al.
Veröffentlicht: (2026)
HTTM: Head-wise Temporal Token Merging for Faster VGGT
von: Wang, Weitian, et al.
Veröffentlicht: (2025)
von: Wang, Weitian, et al.
Veröffentlicht: (2025)
FastVGGT: Training-Free Acceleration of Visual Geometry Transformer
von: Shen, You, et al.
Veröffentlicht: (2025)
von: Shen, You, et al.
Veröffentlicht: (2025)
Reloc-VGGT: Visual Re-localization with Geometry Grounded Transformer
von: Deng, Tianchen, et al.
Veröffentlicht: (2025)
von: Deng, Tianchen, et al.
Veröffentlicht: (2025)
InfiniteVGGT: Visual Geometry Grounded Transformer for Endless Streams
von: Yuan, Shuai, et al.
Veröffentlicht: (2026)
von: Yuan, Shuai, et al.
Veröffentlicht: (2026)
HeSS: Head Sensitivity Score for Sparsity Redistribution in VGGT
von: Kim, Yongsung, et al.
Veröffentlicht: (2026)
von: Kim, Yongsung, et al.
Veröffentlicht: (2026)
UrbanVGGT: Scalable Sidewalk Width Estimation from Street View Images
von: Tan, Kaizhen, et al.
Veröffentlicht: (2026)
von: Tan, Kaizhen, et al.
Veröffentlicht: (2026)
4DLangVGGT: 4D Language-Visual Geometry Grounded Transformer
von: Wu, Xianfeng, et al.
Veröffentlicht: (2025)
von: Wu, Xianfeng, et al.
Veröffentlicht: (2025)
4D-VGGT: A General Foundation Model with SpatioTemporal Awareness for Dynamic Scene Geometry Estimation
von: Wang, Haonan, et al.
Veröffentlicht: (2025)
von: Wang, Haonan, et al.
Veröffentlicht: (2025)
VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold
von: Maggio, Dominic, et al.
Veröffentlicht: (2025)
von: Maggio, Dominic, et al.
Veröffentlicht: (2025)
VGGT-HPE: Reframing Head Pose Estimation as Relative Pose Prediction
von: Vasileiou, Vasiliki, et al.
Veröffentlicht: (2026)
von: Vasileiou, Vasiliki, et al.
Veröffentlicht: (2026)
VGGT-360: Geometry-Consistent Zero-Shot Panoramic Depth Estimation
von: Yuan, Jiayi, et al.
Veröffentlicht: (2026)
von: Yuan, Jiayi, et al.
Veröffentlicht: (2026)
Understanding Attention Mechanism in Video Diffusion Models
von: Liu, Bingyan, et al.
Veröffentlicht: (2025)
von: Liu, Bingyan, et al.
Veröffentlicht: (2025)
Attention Itself Could Retrieve.RetrieveVGGT: Training-Free Long Context Streaming 3D Reconstruction via Query-Key Similarity Retrieval
von: Zou, Zichen, et al.
Veröffentlicht: (2026)
von: Zou, Zichen, et al.
Veröffentlicht: (2026)
StereoVGGT: A Training-Free Visual Geometry Transformer for Stereo Vision
von: Chen, Ziyang, et al.
Veröffentlicht: (2026)
von: Chen, Ziyang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
VGGT-X: When VGGT Meets Dense Novel View Synthesis
von: Liu, Yang, et al.
Veröffentlicht: (2025) -
AVGGT: Rethinking Global Attention for Accelerating VGGT
von: Sun, Xianbing, et al.
Veröffentlicht: (2025) -
VGGT-World: Transforming VGGT into an Autoregressive Geometry World Model
von: Sun, Xiangyu, et al.
Veröffentlicht: (2026) -
VGGT-$Ω$
von: Wang, Jianyuan, et al.
Veröffentlicht: (2026) -
PaceVGGT: Pre-Alternating-Attention Token Pruning for Visual Geometry Transformers
von: Li, Haotang, et al.
Veröffentlicht: (2026)