S-VGGT: Structure-Aware Subscene Decomposition for Scalable 3D Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Xinze, Chen, Pengxu, Wang, Yiyuan, Su, Weifeng, Cheng, Wentao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
QuadBox: Accelerating 3D Gaussian Splatting with Geometry-Aware Boxes
by: Li, Xinze, et al.
Published: (2026)
by: Li, Xinze, et al.
Published: (2026)
4D-VGGT: A General Foundation Model with SpatioTemporal Awareness for Dynamic Scene Geometry Estimation
by: Wang, Haonan, et al.
Published: (2025)
by: Wang, Haonan, et al.
Published: (2025)
VGGT-Occ: Geometry-Grounded and Density-Aware Gated Fusion for 3D Occupancy Prediction
by: Chen, Xun, et al.
Published: (2026)
by: Chen, Xun, et al.
Published: (2026)
VGGT-World: Transforming VGGT into an Autoregressive Geometry World Model
by: Sun, Xiangyu, et al.
Published: (2026)
by: Sun, Xiangyu, et al.
Published: (2026)
VGGT-$Ω$
by: Wang, Jianyuan, et al.
Published: (2026)
by: Wang, Jianyuan, et al.
Published: (2026)
VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection
by: Cao, Yang, et al.
Published: (2026)
by: Cao, Yang, et al.
Published: (2026)
LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token Merging
by: Shu, Zhijian, et al.
Published: (2025)
by: Shu, Zhijian, et al.
Published: (2025)
FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor Attention
by: Wang, Zipeng, et al.
Published: (2025)
by: Wang, Zipeng, et al.
Published: (2025)
VGGT-MPR: VGGT-Enhanced Multimodal Place Recognition in Autonomous Driving Environments
by: Xu, Jingyi, et al.
Published: (2026)
by: Xu, Jingyi, et al.
Published: (2026)
VGGT: Visual Geometry Grounded Transformer
by: Wang, Jianyuan, et al.
Published: (2025)
by: Wang, Jianyuan, et al.
Published: (2025)
VGGT4D: Mining Motion Cues in Visual Geometry Transformers for 4D Scene Reconstruction
by: Hu, Yu, et al.
Published: (2025)
by: Hu, Yu, et al.
Published: (2025)
GPA-VGGT:Adapting VGGT to Large Scale Localization by Self-Supervised Learning with Geometry and Physics Aware Loss
by: Xu, Yangfan, et al.
Published: (2026)
by: Xu, Yangfan, et al.
Published: (2026)
PanoVGGT: Feed-Forward 3D Reconstruction from Panoramic Imagery
by: Guo, Yijing, et al.
Published: (2026)
by: Guo, Yijing, et al.
Published: (2026)
VGGT-X: When VGGT Meets Dense Novel View Synthesis
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
UrbanVGGT: Scalable Sidewalk Width Estimation from Street View Images
by: Tan, Kaizhen, et al.
Published: (2026)
by: Tan, Kaizhen, et al.
Published: (2026)
Leveraging CORAL-Correlation Consistency Network for Semi-Supervised Left Atrium MRI Segmentation
by: Li, Xinze, et al.
Published: (2024)
by: Li, Xinze, et al.
Published: (2024)
VGGT-Edit: Feed-forward Native 3D Scene Editing with Residual Field Prediction
by: Zhu, Kaixin, et al.
Published: (2026)
by: Zhu, Kaixin, et al.
Published: (2026)
HD-VGGT: High-Resolution Visual Geometry Transformer
by: Chen, Tianrun, et al.
Published: (2026)
by: Chen, Tianrun, et al.
Published: (2026)
VGGT-CD: Training-Free Robust Registration for 3D Change Detection
by: Zhang, Wei, et al.
Published: (2026)
by: Zhang, Wei, et al.
Published: (2026)
VGGT-Motion: Motion-Aware Calibration-Free Monocular SLAM for Long-Range Consistency
by: Xiong, Zhuang, et al.
Published: (2026)
by: Xiong, Zhuang, et al.
Published: (2026)
FrameVGGT: Geometry-Aligned Frame-Level Memory for Bounded Streaming VGGT
by: Xu, Zhisong, et al.
Published: (2026)
by: Xu, Zhisong, et al.
Published: (2026)
VGGT-SLAM++
by: Mandal, Avilasha, et al.
Published: (2026)
by: Mandal, Avilasha, et al.
Published: (2026)
PAGE-4D: VGGT-4D Perception via Disentangled Pose and Geometry Estimation
by: Zhou, Kaichen, et al.
Published: (2025)
by: Zhou, Kaichen, et al.
Published: (2025)
SegVGGT: Joint 3D Reconstruction and Instance Segmentation from Multi-View Images
by: Qu, Jinyuan, et al.
Published: (2026)
by: Qu, Jinyuan, et al.
Published: (2026)
Review of Feed-forward 3D Reconstruction: From DUSt3R to VGGT
by: Zhang, Wei, et al.
Published: (2025)
by: Zhang, Wei, et al.
Published: (2025)
4DLangVGGT: 4D Language-Visual Geometry Grounded Transformer
by: Wu, Xianfeng, et al.
Published: (2025)
by: Wu, Xianfeng, et al.
Published: (2025)
Probing the 3D Awareness of Visual Foundation Models
by: Banani, Mohamed El, et al.
Published: (2024)
by: Banani, Mohamed El, et al.
Published: (2024)
StereoVGGT: A Training-Free Visual Geometry Transformer for Stereo Vision
by: Chen, Ziyang, et al.
Published: (2026)
by: Chen, Ziyang, et al.
Published: (2026)
Revisiting 2D Foundation Models for Scalable 3D Medical Image Classification
by: Liu, Han, et al.
Published: (2025)
by: Liu, Han, et al.
Published: (2025)
Building temporally coherent 3D maps with VGGT for memory-efficient Semantic SLAM
by: Dinya, Gergely, et al.
Published: (2025)
by: Dinya, Gergely, et al.
Published: (2025)
DynamicVGGT: Learning Dynamic Point Maps for 4D Scene Reconstruction in Autonomous Driving
by: He, Zhuolin, et al.
Published: (2026)
by: He, Zhuolin, et al.
Published: (2026)
Tracking-Guided 4D Generation: Foundation-Tracker Motion Priors for 3D Model Animation
by: Sun, Su, et al.
Published: (2025)
by: Sun, Su, et al.
Published: (2025)
VGGT-Long: Chunk it, Loop it, Align it -- Pushing VGGT's Limits on Kilometer-scale Long RGB Sequences
by: Deng, Kai, et al.
Published: (2025)
by: Deng, Kai, et al.
Published: (2025)
AVGGT: Rethinking Global Attention for Accelerating VGGT
by: Sun, Xianbing, et al.
Published: (2025)
by: Sun, Xianbing, et al.
Published: (2025)
Reloc-VGGT: Visual Re-localization with Geometry Grounded Transformer
by: Deng, Tianchen, et al.
Published: (2025)
by: Deng, Tianchen, et al.
Published: (2025)
Dense Semantic Matching with VGGT Prior
by: Yang, Songlin, et al.
Published: (2025)
by: Yang, Songlin, et al.
Published: (2025)
An Evaluation of DUSt3R/MASt3R/VGGT 3D Reconstruction on Photogrammetric Aerial Blocks
by: Wu, Xinyi, et al.
Published: (2025)
by: Wu, Xinyi, et al.
Published: (2025)
Structured 3D Latents for Scalable and Versatile 3D Generation
by: Xiang, Jianfeng, et al.
Published: (2024)
by: Xiang, Jianfeng, et al.
Published: (2024)
HCC-3D: Hierarchical Compensatory Compression for 98% 3D Token Reduction in Vision-Language Models
by: Zhang, Liheng, et al.
Published: (2025)
by: Zhang, Liheng, et al.
Published: (2025)
Rank-Aware Agglomeration of Foundation Models for Immunohistochemistry Image Cell Counting
by: Huang, Zuqi, et al.
Published: (2025)
by: Huang, Zuqi, et al.
Published: (2025)
Similar Items
-
QuadBox: Accelerating 3D Gaussian Splatting with Geometry-Aware Boxes
by: Li, Xinze, et al.
Published: (2026) -
4D-VGGT: A General Foundation Model with SpatioTemporal Awareness for Dynamic Scene Geometry Estimation
by: Wang, Haonan, et al.
Published: (2025) -
VGGT-Occ: Geometry-Grounded and Density-Aware Gated Fusion for 3D Occupancy Prediction
by: Chen, Xun, et al.
Published: (2026) -
VGGT-World: Transforming VGGT into an Autoregressive Geometry World Model
by: Sun, Xiangyu, et al.
Published: (2026) -
VGGT-$Ω$
by: Wang, Jianyuan, et al.
Published: (2026)