VGGT: Visual Geometry Grounded Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Jianyuan, Chen, Minghao, Karaev, Nikita, Vedaldi, Andrea, Rupprecht, Christian, Novotny, David |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VGGT-$Ω$
by: Wang, Jianyuan, et al.
Published: (2026)
by: Wang, Jianyuan, et al.
Published: (2026)
CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos
by: Karaev, Nikita, et al.
Published: (2024)
by: Karaev, Nikita, et al.
Published: (2024)
CoTracker: It is Better to Track Together
by: Karaev, Nikita, et al.
Published: (2023)
by: Karaev, Nikita, et al.
Published: (2023)
PoseDiffusion: Solving Pose Estimation via Diffusion-aided Bundle Adjustment
by: Wang, Jianyuan, et al.
Published: (2023)
by: Wang, Jianyuan, et al.
Published: (2023)
Reloc-VGGT: Visual Re-localization with Geometry Grounded Transformer
by: Deng, Tianchen, et al.
Published: (2025)
by: Deng, Tianchen, et al.
Published: (2025)
PartGen: Part-level 3D Generation and Reconstruction with Multi-View Diffusion Models
by: Chen, Minghao, et al.
Published: (2024)
by: Chen, Minghao, et al.
Published: (2024)
InfiniteVGGT: Visual Geometry Grounded Transformer for Endless Streams
by: Yuan, Shuai, et al.
Published: (2026)
by: Yuan, Shuai, et al.
Published: (2026)
OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer
by: Peng, Haosong, et al.
Published: (2025)
by: Peng, Haosong, et al.
Published: (2025)
4DLangVGGT: 4D Language-Visual Geometry Grounded Transformer
by: Wu, Xianfeng, et al.
Published: (2025)
by: Wu, Xianfeng, et al.
Published: (2025)
SHIC: Shape-Image Correspondences with no Keypoint Supervision
by: Shtedritski, Aleksandar, et al.
Published: (2024)
by: Shtedritski, Aleksandar, et al.
Published: (2024)
Splatter Image: Ultra-Fast Single-View 3D Reconstruction
by: Szymanowicz, Stanislaw, et al.
Published: (2023)
by: Szymanowicz, Stanislaw, et al.
Published: (2023)
HD-VGGT: High-Resolution Visual Geometry Transformer
by: Chen, Tianrun, et al.
Published: (2026)
by: Chen, Tianrun, et al.
Published: (2026)
VGGT-World: Transforming VGGT into an Autoregressive Geometry World Model
by: Sun, Xiangyu, et al.
Published: (2026)
by: Sun, Xiangyu, et al.
Published: (2026)
FastVGGT: Training-Free Acceleration of Visual Geometry Transformer
by: Shen, You, et al.
Published: (2025)
by: Shen, You, et al.
Published: (2025)
Invisible Stitch: Generating Smooth 3D Scenes with Depth Inpainting
by: Engstler, Paul, et al.
Published: (2024)
by: Engstler, Paul, et al.
Published: (2024)
Diffusion Models for Open-Vocabulary Segmentation
by: Karazija, Laurynas, et al.
Published: (2023)
by: Karazija, Laurynas, et al.
Published: (2023)
DragAPart: Learning a Part-Level Motion Prior for Articulated Objects
by: Li, Ruining, et al.
Published: (2024)
by: Li, Ruining, et al.
Published: (2024)
SwiftVGGT: A Scalable Visual Geometry Grounded Transformer for Large-Scale Scenes
by: Lee, Jungho, et al.
Published: (2025)
by: Lee, Jungho, et al.
Published: (2025)
FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor Attention
by: Wang, Zipeng, et al.
Published: (2025)
by: Wang, Zipeng, et al.
Published: (2025)
StereoVGGT: A Training-Free Visual Geometry Transformer for Stereo Vision
by: Chen, Ziyang, et al.
Published: (2026)
by: Chen, Ziyang, et al.
Published: (2026)
PaceVGGT: Pre-Alternating-Attention Token Pruning for Visual Geometry Transformers
by: Li, Haotang, et al.
Published: (2026)
by: Li, Haotang, et al.
Published: (2026)
DGE: Direct Gaussian 3D Editing by Consistent Multi-view Editing
by: Chen, Minghao, et al.
Published: (2024)
by: Chen, Minghao, et al.
Published: (2024)
SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotations
by: Wang, Yunnan, et al.
Published: (2026)
by: Wang, Yunnan, et al.
Published: (2026)
Puppet-Master: Scaling Interactive Video Generation as a Motion Prior for Part-Level Dynamics
by: Li, Ruining, et al.
Published: (2024)
by: Li, Ruining, et al.
Published: (2024)
SynCity: Training-Free Generation of 3D Worlds
by: Engstler, Paul, et al.
Published: (2025)
by: Engstler, Paul, et al.
Published: (2025)
Farm3D: Learning Articulated 3D Animals by Distilling 2D Diffusion
by: Jakab, Tomas, et al.
Published: (2023)
by: Jakab, Tomas, et al.
Published: (2023)
DualPM: Dual Posed-Canonical Point Maps for 3D Shape and Pose Reconstruction
by: Kaye, Ben, et al.
Published: (2024)
by: Kaye, Ben, et al.
Published: (2024)
StreamCacheVGGT: Streaming Visual Geometry Transformers with Robust Scoring and Hybrid Cache Compression
by: Liu, Xuanyi, et al.
Published: (2026)
by: Liu, Xuanyi, et al.
Published: (2026)
DriveVGGT: Calibration-Constrained Visual Geometry Transformers for Multi-Camera Autonomous Driving
by: Jia, Xiaosong, et al.
Published: (2025)
by: Jia, Xiaosong, et al.
Published: (2025)
Lightplane: Highly-Scalable Components for Neural 3D Fields
by: Cao, Ang, et al.
Published: (2024)
by: Cao, Ang, et al.
Published: (2024)
TurboVGGT: Fast Visual Geometry Reconstruction with Adaptive Alternating Attention
by: Huang, David, et al.
Published: (2026)
by: Huang, David, et al.
Published: (2026)
VGGT-Occ: Geometry-Grounded and Density-Aware Gated Fusion for 3D Occupancy Prediction
by: Chen, Xun, et al.
Published: (2026)
by: Chen, Xun, et al.
Published: (2026)
Mamba-VGGT: Persistent Long-Sequence Video Geometry Grounded Transformer via External Sliding Window Mamba Memory
by: Deng, Tianchen, et al.
Published: (2026)
by: Deng, Tianchen, et al.
Published: (2026)
XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression
by: Su, Zunhai, et al.
Published: (2026)
by: Su, Zunhai, et al.
Published: (2026)
XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression
by: Su, Zunhai, et al.
Published: (2026)
by: Su, Zunhai, et al.
Published: (2026)
VGGT4D: Mining Motion Cues in Visual Geometry Transformers for 4D Scene Reconstruction
by: Hu, Yu, et al.
Published: (2025)
by: Hu, Yu, et al.
Published: (2025)
Twinner: Shining Light on Digital Twins in a Few Snaps
by: Zarzar, Jesus, et al.
Published: (2025)
by: Zarzar, Jesus, et al.
Published: (2025)
DSO: Aligning 3D Generators with Simulation Feedback for Physical Soundness
by: Li, Ruining, et al.
Published: (2025)
by: Li, Ruining, et al.
Published: (2025)
Learning segmentation from point trajectories
by: Karazija, Laurynas, et al.
Published: (2025)
by: Karazija, Laurynas, et al.
Published: (2025)
VFMF: World Modeling by Forecasting Vision Foundation Model Features
by: Boduljak, Gabrijel, et al.
Published: (2025)
by: Boduljak, Gabrijel, et al.
Published: (2025)
Similar Items
-
VGGT-$Ω$
by: Wang, Jianyuan, et al.
Published: (2026) -
CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos
by: Karaev, Nikita, et al.
Published: (2024) -
CoTracker: It is Better to Track Together
by: Karaev, Nikita, et al.
Published: (2023) -
PoseDiffusion: Solving Pose Estimation via Diffusion-aided Bundle Adjustment
by: Wang, Jianyuan, et al.
Published: (2023) -
Reloc-VGGT: Visual Re-localization with Geometry Grounded Transformer
by: Deng, Tianchen, et al.
Published: (2025)