MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-Experts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Jingnan, Wang, Zhe, Fang, Xianze, Ren, Xingyu, Chen, Zhuo, Liu, Shengqi, Cheng, Yuhao, Lyu, Jiangjing, Yang, Xiaokang, Yan, Yichao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915588456054784
author Gao, Jingnan
Wang, Zhe
Fang, Xianze
Ren, Xingyu
Chen, Zhuo
Liu, Shengqi
Cheng, Yuhao
Lyu, Jiangjing
Yang, Xiaokang
Yan, Yichao
author_facet Gao, Jingnan
Wang, Zhe
Fang, Xianze
Ren, Xingyu
Chen, Zhuo
Liu, Shengqi
Cheng, Yuhao
Lyu, Jiangjing
Yang, Xiaokang
Yan, Yichao
contents Recent advances in language and vision have demonstrated that scaling up model capacity consistently improves performance across diverse tasks. In 3D visual geometry reconstruction, large-scale training has likewise proven effective for learning versatile representations. However, further scaling of 3D models is challenging due to the complexity of geometric supervision and the diversity of 3D data. To overcome these limitations, we propose MoRE, a dense 3D visual foundation model based on a Mixture-of-Experts (MoE) architecture that dynamically routes features to task-specific experts, allowing them to specialize in complementary data aspects and enhance both scalability and adaptability. Aiming to improve robustness under real-world conditions, MoRE incorporates a confidence-based depth refinement module that stabilizes and refines geometric estimation. In addition, it integrates dense semantic features with globally aligned 3D backbone representations for high-fidelity surface normal prediction. MoRE is further optimized with tailored loss functions to ensure robust learning across diverse inputs and multiple geometric tasks. Extensive experiments demonstrate that MoRE achieves state-of-the-art performance across multiple benchmarks and supports effective downstream applications without extra computation.
format Preprint
id arxiv_https___arxiv_org_abs_2510_27234
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-Experts
Gao, Jingnan
Wang, Zhe
Fang, Xianze
Ren, Xingyu
Chen, Zhuo
Liu, Shengqi
Cheng, Yuhao
Lyu, Jiangjing
Yang, Xiaokang
Yan, Yichao
Computer Vision and Pattern Recognition
Recent advances in language and vision have demonstrated that scaling up model capacity consistently improves performance across diverse tasks. In 3D visual geometry reconstruction, large-scale training has likewise proven effective for learning versatile representations. However, further scaling of 3D models is challenging due to the complexity of geometric supervision and the diversity of 3D data. To overcome these limitations, we propose MoRE, a dense 3D visual foundation model based on a Mixture-of-Experts (MoE) architecture that dynamically routes features to task-specific experts, allowing them to specialize in complementary data aspects and enhance both scalability and adaptability. Aiming to improve robustness under real-world conditions, MoRE incorporates a confidence-based depth refinement module that stabilizes and refines geometric estimation. In addition, it integrates dense semantic features with globally aligned 3D backbone representations for high-fidelity surface normal prediction. MoRE is further optimized with tailored loss functions to ensure robust learning across diverse inputs and multiple geometric tasks. Extensive experiments demonstrate that MoRE achieves state-of-the-art performance across multiple benchmarks and supports effective downstream applications without extra computation.
title MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-Experts
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.27234