Unlocking the Power of Critical Factors for 3D Visual Geometry Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Guangkai, Geng, Hua, Zheng, Huanyi, Yin, Songyi, Sun, Yanlong, Chen, Hao, Shen, Chunhua
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913057933885440
author Xu, Guangkai
Geng, Hua
Zheng, Huanyi
Yin, Songyi
Sun, Yanlong
Chen, Hao
Shen, Chunhua
author_facet Xu, Guangkai
Geng, Hua
Zheng, Huanyi
Yin, Songyi
Sun, Yanlong
Chen, Hao
Shen, Chunhua
contents Feed-forward visual geometry estimation has recently made rapid progress. However, an important gap remains: multi-frame models usually produce better cross-frame consistency, yet they often underperform strong per-frame methods on single-frame accuracy. This observation motivates our systematic investigation into the critical factors driving model performance through rigorous ablation studies, which reveals several key insights: 1) Scaling up data diversity and quality unlocks further performance gains even in state-of-the-art visual geometry estimation methods; 2) Commonly adopted confidence-aware loss and gradient-based loss mechanisms may unintentionally hinder performance; 3) Joint supervision through both per-sequence and per-frame alignment improves results, while local region alignment surprisingly degrades performance. Furthermore, we introduce two enhancements to integrate the advantages of optimization-based methods and high-resolution inputs: a consistency loss function that enforces alignment between depth maps, camera parameters, and point maps, and an efficient architectural design that leverages high-resolution information. We integrate these designs into CARVE, a resolution-enhanced model for feed-forward visual geometry estimation. Experiments on point cloud reconstruction, video depth estimation, and camera pose/intrinsic estimation show that CARVE achieves strong and robust performance across diverse benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2604_21713
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Unlocking the Power of Critical Factors for 3D Visual Geometry Estimation
Xu, Guangkai
Geng, Hua
Zheng, Huanyi
Yin, Songyi
Sun, Yanlong
Chen, Hao
Shen, Chunhua
Computer Vision and Pattern Recognition
Feed-forward visual geometry estimation has recently made rapid progress. However, an important gap remains: multi-frame models usually produce better cross-frame consistency, yet they often underperform strong per-frame methods on single-frame accuracy. This observation motivates our systematic investigation into the critical factors driving model performance through rigorous ablation studies, which reveals several key insights: 1) Scaling up data diversity and quality unlocks further performance gains even in state-of-the-art visual geometry estimation methods; 2) Commonly adopted confidence-aware loss and gradient-based loss mechanisms may unintentionally hinder performance; 3) Joint supervision through both per-sequence and per-frame alignment improves results, while local region alignment surprisingly degrades performance. Furthermore, we introduce two enhancements to integrate the advantages of optimization-based methods and high-resolution inputs: a consistency loss function that enforces alignment between depth maps, camera parameters, and point maps, and an efficient architectural design that leverages high-resolution information. We integrate these designs into CARVE, a resolution-enhanced model for feed-forward visual geometry estimation. Experiments on point cloud reconstruction, video depth estimation, and camera pose/intrinsic estimation show that CARVE achieves strong and robust performance across diverse benchmarks.
title Unlocking the Power of Critical Factors for 3D Visual Geometry Estimation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.21713