Unlocking the Power of Critical Factors for 3D Visual Geometry Estimation
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Guangkai, Geng, Hua, Zheng, Huanyi, Yin, Songyi, Sun, Yanlong, Chen, Hao, Shen, Chunhua |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GeoBench: Benchmarking and Analyzing Monocular Geometry Estimation Models
by: Ge, Yongtao, et al.
Published: (2024)
by: Ge, Yongtao, et al.
Published: (2024)
DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks
by: Zhao, Canyu, et al.
Published: (2025)
by: Zhao, Canyu, et al.
Published: (2025)
POMATO: Marrying Pointmap Matching with Temporal Motion for Dynamic 3D Reconstruction
by: Zhang, Songyan, et al.
Published: (2025)
by: Zhang, Songyan, et al.
Published: (2025)
What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?
by: Xu, Guangkai, et al.
Published: (2024)
by: Xu, Guangkai, et al.
Published: (2024)
Unleashing the Potential of the Diffusion Model in Few-shot Semantic Segmentation
by: Zhu, Muzhi, et al.
Published: (2024)
by: Zhu, Muzhi, et al.
Published: (2024)
Generative Video Matting
by: Ge, Yongtao, et al.
Published: (2025)
by: Ge, Yongtao, et al.
Published: (2025)
$π^3$: Permutation-Equivariant Visual Geometry Learning
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
Pixel-Perfect Visual Geometry Estimation
by: Xu, Gangwei, et al.
Published: (2026)
by: Xu, Gangwei, et al.
Published: (2026)
CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models
by: Huang, Junming, et al.
Published: (2026)
by: Huang, Junming, et al.
Published: (2026)
GeoHand: Unlocking Prior Geometry Knowledge for Monocular 3D Hand Reconstruction
by: Lin, Weiquan, et al.
Published: (2026)
by: Lin, Weiquan, et al.
Published: (2026)
GeoMotion: Rethinking Motion Segmentation via Latent 4D Geometry
by: He, Xiankang, et al.
Published: (2026)
by: He, Xiankang, et al.
Published: (2026)
TC-AE: Unlocking Token Capacity for Deep Compression Autoencoders
by: Li, Teng, et al.
Published: (2026)
by: Li, Teng, et al.
Published: (2026)
Exploring Spatial Intelligence from a Generative Perspective
by: Zhu, Muzhi, et al.
Published: (2026)
by: Zhu, Muzhi, et al.
Published: (2026)
4DVGGT-D: 4D Visual Geometry Transformer with Improved Dynamic Depth Estimation
by: Zang, Ying, et al.
Published: (2026)
by: Zang, Ying, et al.
Published: (2026)
Reg3D: Reconstructive Geometry Instruction Tuning for 3D Scene Understanding
by: Zheng, Hongpei, et al.
Published: (2025)
by: Zheng, Hongpei, et al.
Published: (2025)
GeoWizard: Unleashing the Diffusion Priors for 3D Geometry Estimation from a Single Image
by: Fu, Xiao, et al.
Published: (2024)
by: Fu, Xiao, et al.
Published: (2024)
Metric3Dv2: A Versatile Monocular Geometric Foundation Model for Zero-shot Metric Depth and Surface Normal Estimation
by: Hu, Mu, et al.
Published: (2024)
by: Hu, Mu, et al.
Published: (2024)
Reflections Unlock: Geometry-Aware Reflection Disentanglement in 3D Gaussian Splatting for Photorealistic Scenes Rendering
by: Song, Jiayi, et al.
Published: (2025)
by: Song, Jiayi, et al.
Published: (2025)
GaussianGrow: Geometry-aware Gaussian Growing from 3D Point Clouds with Text Guidance
by: Zhang, Weiqi, et al.
Published: (2026)
by: Zhang, Weiqi, et al.
Published: (2026)
Geo-Align: Video Generation Alignment via Metric Geometry Reward
by: Li, Zizun, et al.
Published: (2026)
by: Li, Zizun, et al.
Published: (2026)
UniPR-3D: Towards Universal Visual Place Recognition with Visual Geometry Grounded Transformer
by: Deng, Tianchen, et al.
Published: (2025)
by: Deng, Tianchen, et al.
Published: (2025)
MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision
by: Wang, Ruicheng, et al.
Published: (2024)
by: Wang, Ruicheng, et al.
Published: (2024)
One View, Many Worlds: Single-Image to 3D Object Meets Generative Domain Randomization for One-Shot 6D Pose Estimation
by: Geng, Zheng, et al.
Published: (2025)
by: Geng, Zheng, et al.
Published: (2025)
Lifting the Veil on Visual Information Flow in MLLMs: Unlocking Pathways to Faster Inference
by: Yin, Hao, et al.
Published: (2025)
by: Yin, Hao, et al.
Published: (2025)
TrajVG: 3D Trajectory-Coupled Visual Geometry Learning
by: Miao, Xingyu, et al.
Published: (2026)
by: Miao, Xingyu, et al.
Published: (2026)
DiffCalib: Reformulating Monocular Camera Calibration as Diffusion-Based Dense Incident Map Generation
by: He, Xiankang, et al.
Published: (2024)
by: He, Xiankang, et al.
Published: (2024)
VIMCAN: Visual-Inertial 3D Human Pose Estimation with Hybrid Mamba-Cross-Attention Network
by: Yang, Zepeng, et al.
Published: (2026)
by: Yang, Zepeng, et al.
Published: (2026)
Language as an Anchor: Preserving Relative Visual Geometry for Domain Incremental Learning
by: Geng, Shuyi, et al.
Published: (2025)
by: Geng, Shuyi, et al.
Published: (2025)
Tinker: Diffusion's Gift to 3D--Multi-View Consistent Editing From Sparse Inputs without Per-Scene Optimization
by: Zhao, Canyu, et al.
Published: (2025)
by: Zhao, Canyu, et al.
Published: (2025)
Unlocking 3D Affordance Segmentation with 2D Semantic Knowledge
by: Huang, Yu, et al.
Published: (2025)
by: Huang, Yu, et al.
Published: (2025)
RGM: A Robust Generalizable Matching Model
by: Zhang, Songyan, et al.
Published: (2023)
by: Zhang, Songyan, et al.
Published: (2023)
Interp3R: Continuous-time 3D Geometry Estimation with Frames and Events
by: Guo, Shuang, et al.
Published: (2026)
by: Guo, Shuang, et al.
Published: (2026)
Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation
by: Bai, Weimin, et al.
Published: (2025)
by: Bai, Weimin, et al.
Published: (2025)
GeoWorld: Unlocking the Potential of Geometry Models to Facilitate High-Fidelity 3D Scene Generation
by: Wan, Yuhao, et al.
Published: (2025)
by: Wan, Yuhao, et al.
Published: (2025)
Flow3r: Factored Flow Prediction for Scalable Visual Geometry Learning
by: Cong, Zhongxiao, et al.
Published: (2026)
by: Cong, Zhongxiao, et al.
Published: (2026)
FastVGGT: Training-Free Acceleration of Visual Geometry Transformer
by: Shen, You, et al.
Published: (2025)
by: Shen, You, et al.
Published: (2025)
LongStream: Long-Sequence Streaming Autoregressive Visual Geometry
by: Cheng, Chong, et al.
Published: (2026)
by: Cheng, Chong, et al.
Published: (2026)
Scaling up Multi-domain Semantic Segmentation with Sentence Embeddings
by: Yin, Wei, et al.
Published: (2022)
by: Yin, Wei, et al.
Published: (2022)
SurfSurg6D: Geometry Consistent Dense Correspondence for Textureless Surgical Instrument Pose Estimation
by: Shen, Daiyun, et al.
Published: (2026)
by: Shen, Daiyun, et al.
Published: (2026)
MovieDreamer: Hierarchical Generation for Coherent Long Visual Sequence
by: Zhao, Canyu, et al.
Published: (2024)
by: Zhao, Canyu, et al.
Published: (2024)
Similar Items
-
GeoBench: Benchmarking and Analyzing Monocular Geometry Estimation Models
by: Ge, Yongtao, et al.
Published: (2024) -
DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks
by: Zhao, Canyu, et al.
Published: (2025) -
POMATO: Marrying Pointmap Matching with Temporal Motion for Dynamic 3D Reconstruction
by: Zhang, Songyan, et al.
Published: (2025) -
What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?
by: Xu, Guangkai, et al.
Published: (2024) -
Unleashing the Potential of the Diffusion Model in Few-shot Semantic Segmentation
by: Zhu, Muzhi, et al.
Published: (2024)