GPA-VGGT:Adapting VGGT to Large Scale Localization by Self-Supervised Learning with Geometry and Physics Aware Loss
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Yangfan, Zhang, Lilian, He, Xiaofeng, Wu, Pengdong, Wu, Wenqi, Mao, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SLAM in the Dark: Self-Supervised Learning of Pose, Depth and Loop-Closure from Thermal Images
by: Xu, Yangfan, et al.
Published: (2025)
by: Xu, Yangfan, et al.
Published: (2025)
VGGT-SLAM++
by: Mandal, Avilasha, et al.
Published: (2026)
by: Mandal, Avilasha, et al.
Published: (2026)
SceneVGGT: VGGT-based online 3D semantic SLAM for indoor scene understanding and navigation
by: Gelencsér-Horváth, Anna, et al.
Published: (2026)
by: Gelencsér-Horváth, Anna, et al.
Published: (2026)
FrameVGGT: Geometry-Aligned Frame-Level Memory for Bounded Streaming VGGT
by: Xu, Zhisong, et al.
Published: (2026)
by: Xu, Zhisong, et al.
Published: (2026)
VGGT-DP: Generalizable Robot Control via Vision Foundation Models
by: Ge, Shijia, et al.
Published: (2025)
by: Ge, Shijia, et al.
Published: (2025)
VGGT-World: Transforming VGGT into an Autoregressive Geometry World Model
by: Sun, Xiangyu, et al.
Published: (2026)
by: Sun, Xiangyu, et al.
Published: (2026)
LiDAR-VGGT: Cross-Modal Coarse-to-Fine Fusion for Globally Consistent and Metric-Scale Dense Mapping
by: Wang, Lijie, et al.
Published: (2025)
by: Wang, Lijie, et al.
Published: (2025)
HyVGGT-VO: Tightly Coupled Hybrid Dense Visual Odometry with Feed-Forward Models
by: Pan, Junxiang, et al.
Published: (2026)
by: Pan, Junxiang, et al.
Published: (2026)
VGGT-SLAM 2.0: Real-time Dense Feed-forward Scene Reconstruction
by: Maggio, Dominic, et al.
Published: (2026)
by: Maggio, Dominic, et al.
Published: (2026)
VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection
by: Cao, Yang, et al.
Published: (2026)
by: Cao, Yang, et al.
Published: (2026)
LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token Merging
by: Shu, Zhijian, et al.
Published: (2025)
by: Shu, Zhijian, et al.
Published: (2025)
3D-Mix for VLA: A Plug-and-Play Module for Integrating VGGT-based 3D Information into Vision-Language-Action Models
by: Yu, Bin, et al.
Published: (2026)
by: Yu, Bin, et al.
Published: (2026)
VGGT: Visual Geometry Grounded Transformer
by: Wang, Jianyuan, et al.
Published: (2025)
by: Wang, Jianyuan, et al.
Published: (2025)
VGGT-Segmentor: Geometry-Enhanced Cross-View Segmentation
by: Gao, Yulu, et al.
Published: (2026)
by: Gao, Yulu, et al.
Published: (2026)
Reloc-VGGT: Visual Re-localization with Geometry Grounded Transformer
by: Deng, Tianchen, et al.
Published: (2025)
by: Deng, Tianchen, et al.
Published: (2025)
VGC-RIO: A Tightly Integrated Radar-Inertial Odometry with Spatial Weighted Doppler Velocity and Local Geometric Constrained RCS Histograms
by: Xiang, Jianguang, et al.
Published: (2025)
by: Xiang, Jianguang, et al.
Published: (2025)
VGGT-$Ω$
by: Wang, Jianyuan, et al.
Published: (2026)
by: Wang, Jianyuan, et al.
Published: (2026)
VGGT-MPR: VGGT-Enhanced Multimodal Place Recognition in Autonomous Driving Environments
by: Xu, Jingyi, et al.
Published: (2026)
by: Xu, Jingyi, et al.
Published: (2026)
SwiftVGGT: A Scalable Visual Geometry Grounded Transformer for Large-Scale Scenes
by: Lee, Jungho, et al.
Published: (2025)
by: Lee, Jungho, et al.
Published: (2025)
HD-VGGT: High-Resolution Visual Geometry Transformer
by: Chen, Tianrun, et al.
Published: (2026)
by: Chen, Tianrun, et al.
Published: (2026)
VGGT-X: When VGGT Meets Dense Novel View Synthesis
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
TurboVGGT: Fast Visual Geometry Reconstruction with Adaptive Alternating Attention
by: Huang, David, et al.
Published: (2026)
by: Huang, David, et al.
Published: (2026)
GPA-RAM: Grasp-Pretraining Augmented Robotic Attention Mamba for Spatial Task Learning
by: Sheng, Juyi, et al.
Published: (2025)
by: Sheng, Juyi, et al.
Published: (2025)
4DLangVGGT: 4D Language-Visual Geometry Grounded Transformer
by: Wu, Xianfeng, et al.
Published: (2025)
by: Wu, Xianfeng, et al.
Published: (2025)
SP-VINS: A Hybrid Stereo Visual Inertial Navigation System based on Implicit Environmental Map
by: Du, Xueyu, et al.
Published: (2025)
by: Du, Xueyu, et al.
Published: (2025)
VGGT-Long: Chunk it, Loop it, Align it -- Pushing VGGT's Limits on Kilometer-scale Long RGB Sequences
by: Deng, Kai, et al.
Published: (2025)
by: Deng, Kai, et al.
Published: (2025)
PO-MSCKF: An Efficient Visual-Inertial Odometry by Reconstructing the Multi-State Constrained Kalman Filter with the Pose-only Theory
by: Du, Xueyu, et al.
Published: (2024)
by: Du, Xueyu, et al.
Published: (2024)
InfiniteVGGT: Visual Geometry Grounded Transformer for Endless Streams
by: Yuan, Shuai, et al.
Published: (2026)
by: Yuan, Shuai, et al.
Published: (2026)
FastVGGT: Training-Free Acceleration of Visual Geometry Transformer
by: Shen, You, et al.
Published: (2025)
by: Shen, You, et al.
Published: (2025)
Information-Theoretic Geometry Optimization and Physics-Aware Learning for Calibration-Free Magnetic Localization
by: Xie, Wenxuan, et al.
Published: (2026)
by: Xie, Wenxuan, et al.
Published: (2026)
FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor Attention
by: Wang, Zipeng, et al.
Published: (2025)
by: Wang, Zipeng, et al.
Published: (2025)
SP-VIO: Robust and Efficient Filter-Based Visual Inertial Odometry with State Transformation Model and Pose-Only Visual Description
by: Du, Xueyu, et al.
Published: (2024)
by: Du, Xueyu, et al.
Published: (2024)
OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer
by: Peng, Haosong, et al.
Published: (2025)
by: Peng, Haosong, et al.
Published: (2025)
PaceVGGT: Pre-Alternating-Attention Token Pruning for Visual Geometry Transformers
by: Li, Haotang, et al.
Published: (2026)
by: Li, Haotang, et al.
Published: (2026)
VGGT-Occ: Geometry-Grounded and Density-Aware Gated Fusion for 3D Occupancy Prediction
by: Chen, Xun, et al.
Published: (2026)
by: Chen, Xun, et al.
Published: (2026)
VGGT-360: Geometry-Consistent Zero-Shot Panoramic Depth Estimation
by: Yuan, Jiayi, et al.
Published: (2026)
by: Yuan, Jiayi, et al.
Published: (2026)
AVGGT: Rethinking Global Attention for Accelerating VGGT
by: Sun, Xianbing, et al.
Published: (2025)
by: Sun, Xianbing, et al.
Published: (2025)
Adapting Skills to Novel Grasps: A Self-Supervised Approach
by: Papagiannis, Georgios, et al.
Published: (2024)
by: Papagiannis, Georgios, et al.
Published: (2024)
Dense Semantic Matching with VGGT Prior
by: Yang, Songlin, et al.
Published: (2025)
by: Yang, Songlin, et al.
Published: (2025)
PC-Planner: Physics-Constrained Self-Supervised Learning for Robust Neural Motion Planning with Shape-Aware Distance Function
by: Shen, Xujie, et al.
Published: (2024)
by: Shen, Xujie, et al.
Published: (2024)
Similar Items
-
SLAM in the Dark: Self-Supervised Learning of Pose, Depth and Loop-Closure from Thermal Images
by: Xu, Yangfan, et al.
Published: (2025) -
VGGT-SLAM++
by: Mandal, Avilasha, et al.
Published: (2026) -
SceneVGGT: VGGT-based online 3D semantic SLAM for indoor scene understanding and navigation
by: Gelencsér-Horváth, Anna, et al.
Published: (2026) -
FrameVGGT: Geometry-Aligned Frame-Level Memory for Bounded Streaming VGGT
by: Xu, Zhisong, et al.
Published: (2026) -
VGGT-DP: Generalizable Robot Control via Vision Foundation Models
by: Ge, Shijia, et al.
Published: (2025)