AVGGT: Rethinking Global Attention for Accelerating VGGT
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Xianbing, Zhu, Zhikai, Lou, Zhengyu, Yang, Bo, Tang, Jinyang, Zhang, Liqing, Wang, He, Zhang, Jianfu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DirectTryOn: One-Step Virtual Try-On via Straightened Conditional Transport
by: Sun, Xianbing, et al.
Published: (2026)
by: Sun, Xianbing, et al.
Published: (2026)
Towards Source-Aware Object Swapping with Initial Noise Perturbation
by: Zhan, Jiahui, et al.
Published: (2026)
by: Zhan, Jiahui, et al.
Published: (2026)
DS-VTON: An Enhanced Dual-Scale Coarse-to-Fine Framework for Virtual Try-On
by: Sun, Xianbing, et al.
Published: (2025)
by: Sun, Xianbing, et al.
Published: (2025)
FW-VTON: Flattening-and-Warping for Person-to-Person Virtual Try-on
by: Wang, Zheng, et al.
Published: (2025)
by: Wang, Zheng, et al.
Published: (2025)
Enhancing Domain Generalization in 3D Human Pose Estimation through Controllable Generative Augmentation
by: Hu, Xinhao, et al.
Published: (2026)
by: Hu, Xinhao, et al.
Published: (2026)
VGGT-X: When VGGT Meets Dense Novel View Synthesis
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
VGGT-World: Transforming VGGT into an Autoregressive Geometry World Model
by: Sun, Xiangyu, et al.
Published: (2026)
by: Sun, Xiangyu, et al.
Published: (2026)
VTONGuard: Automatic Detection and Authentication of AI-Generated Virtual Try-On Content
by: Wu, Shengyi, et al.
Published: (2026)
by: Wu, Shengyi, et al.
Published: (2026)
FastVGGT: Training-Free Acceleration of Visual Geometry Transformer
by: Shen, You, et al.
Published: (2025)
by: Shen, You, et al.
Published: (2025)
PaceVGGT: Pre-Alternating-Attention Token Pruning for Visual Geometry Transformers
by: Li, Haotang, et al.
Published: (2026)
by: Li, Haotang, et al.
Published: (2026)
DomainGallery: Few-shot Domain-driven Image Generation by Attribute-centric Finetuning
by: Duan, Yuxuan, et al.
Published: (2024)
by: Duan, Yuxuan, et al.
Published: (2024)
Any3DAvatar: Fast and High-Quality Full-Head 3D Avatar Reconstruction from Single Portrait Image
by: Gao, Yujie, et al.
Published: (2026)
by: Gao, Yujie, et al.
Published: (2026)
VGGT-$Ω$
by: Wang, Jianyuan, et al.
Published: (2026)
by: Wang, Jianyuan, et al.
Published: (2026)
Assessing Image Inpainting via Re-Inpainting Self-Consistency Evaluation
by: Chen, Tianyi, et al.
Published: (2024)
by: Chen, Tianyi, et al.
Published: (2024)
High-Quality 3D Head Reconstruction from Any Single Portrait Image
by: Zhang, Jianfu, et al.
Published: (2025)
by: Zhang, Jianfu, et al.
Published: (2025)
Robustness in AI-Generated Detection: Enhancing Resistance to Adversarial Attacks
by: Haoxuan, Sun, et al.
Published: (2025)
by: Haoxuan, Sun, et al.
Published: (2025)
User-Friendly Customized Generation with Multi-Modal Prompts
by: Zhong, Linhao, et al.
Published: (2024)
by: Zhong, Linhao, et al.
Published: (2024)
InfiniteVGGT: Visual Geometry Grounded Transformer for Endless Streams
by: Yuan, Shuai, et al.
Published: (2026)
by: Yuan, Shuai, et al.
Published: (2026)
Self-Supervised Vision Transformer for Enhanced Virtual Clothes Try-On
by: Lu, Lingxiao, et al.
Published: (2024)
by: Lu, Lingxiao, et al.
Published: (2024)
VGGT-Segmentor: Geometry-Enhanced Cross-View Segmentation
by: Gao, Yulu, et al.
Published: (2026)
by: Gao, Yulu, et al.
Published: (2026)
FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor Attention
by: Wang, Zipeng, et al.
Published: (2025)
by: Wang, Zipeng, et al.
Published: (2025)
GPA-VGGT:Adapting VGGT to Large Scale Localization by Self-Supervised Learning with Geometry and Physics Aware Loss
by: Xu, Yangfan, et al.
Published: (2026)
by: Xu, Yangfan, et al.
Published: (2026)
TurboVGGT: Fast Visual Geometry Reconstruction with Adaptive Alternating Attention
by: Huang, David, et al.
Published: (2026)
by: Huang, David, et al.
Published: (2026)
Analyzing the Mechanism of Attention Collapse in VGGT from a Dynamics Perspective
by: Li, Huan, et al.
Published: (2025)
by: Li, Huan, et al.
Published: (2025)
WildFake: A Large-scale Challenging Dataset for AI-Generated Images Detection
by: Hong, Yan, et al.
Published: (2024)
by: Hong, Yan, et al.
Published: (2024)
ComFusion: Personalized Subject Generation in Multiple Specific Scenes From Single Image
by: Hong, Yan, et al.
Published: (2024)
by: Hong, Yan, et al.
Published: (2024)
VGGT-Long: Chunk it, Loop it, Align it -- Pushing VGGT's Limits on Kilometer-scale Long RGB Sequences
by: Deng, Kai, et al.
Published: (2025)
by: Deng, Kai, et al.
Published: (2025)
HD-VGGT: High-Resolution Visual Geometry Transformer
by: Chen, Tianrun, et al.
Published: (2026)
by: Chen, Tianrun, et al.
Published: (2026)
VGGT-MPR: VGGT-Enhanced Multimodal Place Recognition in Autonomous Driving Environments
by: Xu, Jingyi, et al.
Published: (2026)
by: Xu, Jingyi, et al.
Published: (2026)
FrameVGGT: Geometry-Aligned Frame-Level Memory for Bounded Streaming VGGT
by: Xu, Zhisong, et al.
Published: (2026)
by: Xu, Zhisong, et al.
Published: (2026)
VGGT-SLAM++
by: Mandal, Avilasha, et al.
Published: (2026)
by: Mandal, Avilasha, et al.
Published: (2026)
Dense Semantic Matching with VGGT Prior
by: Yang, Songlin, et al.
Published: (2025)
by: Yang, Songlin, et al.
Published: (2025)
Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs
by: Ji, Yikun, et al.
Published: (2025)
by: Ji, Yikun, et al.
Published: (2025)
LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token Merging
by: Shu, Zhijian, et al.
Published: (2025)
by: Shu, Zhijian, et al.
Published: (2025)
VGGT: Visual Geometry Grounded Transformer
by: Wang, Jianyuan, et al.
Published: (2025)
by: Wang, Jianyuan, et al.
Published: (2025)
Reloc-VGGT: Visual Re-localization with Geometry Grounded Transformer
by: Deng, Tianchen, et al.
Published: (2025)
by: Deng, Tianchen, et al.
Published: (2025)
Rethinking Classifier Re-Training in Long-Tailed Recognition: A Simple Logits Retargeting Approach
by: Lu, Han, et al.
Published: (2024)
by: Lu, Han, et al.
Published: (2024)
VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection
by: Cao, Yang, et al.
Published: (2026)
by: Cao, Yang, et al.
Published: (2026)
DynamicVGGT: Learning Dynamic Point Maps for 4D Scene Reconstruction in Autonomous Driving
by: He, Zhuolin, et al.
Published: (2026)
by: He, Zhuolin, et al.
Published: (2026)
UrbanVGGT: Scalable Sidewalk Width Estimation from Street View Images
by: Tan, Kaizhen, et al.
Published: (2026)
by: Tan, Kaizhen, et al.
Published: (2026)
Similar Items
-
DirectTryOn: One-Step Virtual Try-On via Straightened Conditional Transport
by: Sun, Xianbing, et al.
Published: (2026) -
Towards Source-Aware Object Swapping with Initial Noise Perturbation
by: Zhan, Jiahui, et al.
Published: (2026) -
DS-VTON: An Enhanced Dual-Scale Coarse-to-Fine Framework for Virtual Try-On
by: Sun, Xianbing, et al.
Published: (2025) -
FW-VTON: Flattening-and-Warping for Person-to-Person Virtual Try-on
by: Wang, Zheng, et al.
Published: (2025) -
Enhancing Domain Generalization in 3D Human Pose Estimation through Controllable Generative Augmentation
by: Hu, Xinhao, et al.
Published: (2026)