Saved in:
| Main Authors: | Deng, Tianchen, Chen, Xun, Li, Ziming, Shen, Hongming, Wang, Danwei, Civera, Javier, Wang, Hesheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.21078 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VGGT-Occ: Geometry-Grounded and Density-Aware Gated Fusion for 3D Occupancy Prediction
by: Chen, Xun, et al.
Published: (2026)
by: Chen, Xun, et al.
Published: (2026)
Reloc-VGGT: Visual Re-localization with Geometry Grounded Transformer
by: Deng, Tianchen, et al.
Published: (2025)
by: Deng, Tianchen, et al.
Published: (2025)
UniLGL: Learning Uniform Place Recognition for FOV-limited/Panoramic LiDAR Global Localization
by: Shen, Hongming, et al.
Published: (2025)
by: Shen, Hongming, et al.
Published: (2025)
Optimal Transport Aggregation for Visual Place Recognition
by: Izquierdo, Sergio, et al.
Published: (2023)
by: Izquierdo, Sergio, et al.
Published: (2023)
Compact 3D Gaussian Splatting For Dense Visual SLAM
by: Deng, Tianchen, et al.
Published: (2024)
by: Deng, Tianchen, et al.
Published: (2024)
Close, But Not There: Boosting Geographic Distance Sensitivity in Visual Place Recognition
by: Izquierdo, Sergio, et al.
Published: (2024)
by: Izquierdo, Sergio, et al.
Published: (2024)
VPGS-SLAM: Voxel-based Progressive 3D Gaussian SLAM in Large-Scale Scenes
by: Deng, Tianchen, et al.
Published: (2025)
by: Deng, Tianchen, et al.
Published: (2025)
MCN-SLAM: Multi-Agent Collaborative Neural SLAM with Hybrid Implicit Neural Scene Representation
by: Deng, Tianchen, et al.
Published: (2025)
by: Deng, Tianchen, et al.
Published: (2025)
What Is The Best 3D Scene Representation for Robotics? From Geometric to Foundation Models
by: Deng, Tianchen, et al.
Published: (2025)
by: Deng, Tianchen, et al.
Published: (2025)
Mamba-VGGT: Persistent Long-Sequence Video Geometry Grounded Transformer via External Sliding Window Mamba Memory
by: Deng, Tianchen, et al.
Published: (2026)
by: Deng, Tianchen, et al.
Published: (2026)
UniPR: Unified Object-level Real-to-Sim Perception and Reconstruction from a Single Stereo Pair
by: Zhang, Chuanrui, et al.
Published: (2026)
by: Zhang, Chuanrui, et al.
Published: (2026)
NeSLAM: Neural Implicit Mapping and Self-Supervised Feature Tracking With Depth Completion and Denoising
by: Deng, Tianchen, et al.
Published: (2024)
by: Deng, Tianchen, et al.
Published: (2024)
DC-VLAQ: Query-Residual Aggregation for Robust Visual Place Recognition
by: Zhu, Hanyu, et al.
Published: (2026)
by: Zhu, Hanyu, et al.
Published: (2026)
BEV$^2$PR: BEV-Enhanced Visual Place Recognition with Structural Cues
by: Ge, Fudong, et al.
Published: (2024)
by: Ge, Fudong, et al.
Published: (2024)
Incremental Joint Learning of Depth, Pose and Implicit Scene Representation on Monocular Camera in Large-scale Scenes
by: Deng, Tianchen, et al.
Published: (2024)
by: Deng, Tianchen, et al.
Published: (2024)
MUT3R: Motion-aware Updating Transformer for Dynamic 3D Reconstruction
by: Shen, Guole, et al.
Published: (2025)
by: Shen, Guole, et al.
Published: (2025)
UniGround: Universal 3D Visual Grounding via Training-Free Scene Parsing
by: Zhang, Jiaxi, et al.
Published: (2026)
by: Zhang, Jiaxi, et al.
Published: (2026)
Keep It CALM: Toward Calibration-Free Kilometer-Level SLAM with Visual Geometry Foundation Models via an Assistant Eye
by: Zhang, Tianjun, et al.
Published: (2026)
by: Zhang, Tianjun, et al.
Published: (2026)
MAG-VLAQ: Multi-modal Aerial-Ground Query Aggregation for Cross-View Place Recognition
by: Xu, Zhengyi, et al.
Published: (2026)
by: Xu, Zhengyi, et al.
Published: (2026)
Towards the Vision-Sound-Language-Action Paradigm: The HEAR Framework for Sound-Centric Manipulation
by: Nie, Chang, et al.
Published: (2026)
by: Nie, Chang, et al.
Published: (2026)
VGGT: Visual Geometry Grounded Transformer
by: Wang, Jianyuan, et al.
Published: (2025)
by: Wang, Jianyuan, et al.
Published: (2025)
Look Ma, No Ground Truth! Ground-Truth-Free Tuning of Structure from Motion and Visual SLAM
by: Fontan, Alejandro, et al.
Published: (2024)
by: Fontan, Alejandro, et al.
Published: (2024)
RoGs: Large Scale Road Surface Reconstruction with Meshgrid Gaussian
by: Feng, Zhiheng, et al.
Published: (2024)
by: Feng, Zhiheng, et al.
Published: (2024)
MID: A Self-supervised Multimodal Iterative Denoising Framework
by: Nie, Chang, et al.
Published: (2025)
by: Nie, Chang, et al.
Published: (2025)
Uni-Animator: Towards Unified Visual Colorization
by: Chen, Xinyuan, et al.
Published: (2026)
by: Chen, Xinyuan, et al.
Published: (2026)
DefVINS: Visual-Inertial Odometry for Deformable Scenes
by: Cerezo, Samuel, et al.
Published: (2026)
by: Cerezo, Samuel, et al.
Published: (2026)
UniLoc: Towards Universal Place Recognition Using Any Single Modality
by: Xia, Yan, et al.
Published: (2024)
by: Xia, Yan, et al.
Published: (2024)
S$^2$-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural Guidance
by: Xu, Beining, et al.
Published: (2025)
by: Xu, Beining, et al.
Published: (2025)
QVGGT: Post-Training Quantized Visual Geometry Grounded Transformer
by: Pan, Zhizhen, et al.
Published: (2026)
by: Pan, Zhizhen, et al.
Published: (2026)
Quantized Visual Geometry Grounded Transformer
by: Feng, Weilun, et al.
Published: (2025)
by: Feng, Weilun, et al.
Published: (2025)
Feature Complementation Architecture for Visual Place Recognition
by: Wang, Weiwei, et al.
Published: (2025)
by: Wang, Weiwei, et al.
Published: (2025)
EDTformer: An Efficient Decoder Transformer for Visual Place Recognition
by: Jin, Tong, et al.
Published: (2024)
by: Jin, Tong, et al.
Published: (2024)
VDNA-PR: Using General Dataset Representations for Robust Sequential Visual Place Recognition
by: Ramtoula, Benjamin, et al.
Published: (2024)
by: Ramtoula, Benjamin, et al.
Published: (2024)
Regressing Transformers for Data-efficient Visual Place Recognition
by: Leyva-Vallina, María, et al.
Published: (2024)
by: Leyva-Vallina, María, et al.
Published: (2024)
4DLangVGGT: 4D Language-Visual Geometry Grounded Transformer
by: Wu, Xianfeng, et al.
Published: (2025)
by: Wu, Xianfeng, et al.
Published: (2025)
DIAL-GS: Dynamic Instance Aware Reconstruction for Label-free Street Scenes with 4D Gaussian Splatting
by: Su, Chenpeng, et al.
Published: (2025)
by: Su, Chenpeng, et al.
Published: (2025)
MVGGT: Multimodal Visual Geometry Grounded Transformer for Multiview 3D Referring Expression Segmentation
by: Wu, Changli, et al.
Published: (2026)
by: Wu, Changli, et al.
Published: (2026)
PLGSLAM: Progressive Neural Scene Represenation with Local to Global Bundle Adjustment
by: Deng, Tianchen, et al.
Published: (2023)
by: Deng, Tianchen, et al.
Published: (2023)
Towards Seamless Adaptation of Pre-trained Models for Visual Place Recognition
by: Lu, Feng, et al.
Published: (2024)
by: Lu, Feng, et al.
Published: (2024)
VG3T: Visual Geometry Grounded Gaussian Transformer
by: Kim, Junho, et al.
Published: (2025)
by: Kim, Junho, et al.
Published: (2025)
Similar Items
-
VGGT-Occ: Geometry-Grounded and Density-Aware Gated Fusion for 3D Occupancy Prediction
by: Chen, Xun, et al.
Published: (2026) -
Reloc-VGGT: Visual Re-localization with Geometry Grounded Transformer
by: Deng, Tianchen, et al.
Published: (2025) -
UniLGL: Learning Uniform Place Recognition for FOV-limited/Panoramic LiDAR Global Localization
by: Shen, Hongming, et al.
Published: (2025) -
Optimal Transport Aggregation for Visual Place Recognition
by: Izquierdo, Sergio, et al.
Published: (2023) -
Compact 3D Gaussian Splatting For Dense Visual SLAM
by: Deng, Tianchen, et al.
Published: (2024)