DiMeR: Disentangled Mesh Reconstruction Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Lutao, Lin, Jiantao, Chen, Kanghao, Ge, Wenhang, Yang, Xin, Jiang, Yifan, Lyu, Yuanhuiyi, Zheng, Xu, Li, Yinchuan, Chen, Yingcong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913859234693120
author Jiang, Lutao
Lin, Jiantao
Chen, Kanghao
Ge, Wenhang
Yang, Xin
Jiang, Yifan
Lyu, Yuanhuiyi
Zheng, Xu
Li, Yinchuan
Chen, Yingcong
author_facet Jiang, Lutao
Lin, Jiantao
Chen, Kanghao
Ge, Wenhang
Yang, Xin
Jiang, Yifan
Lyu, Yuanhuiyi
Zheng, Xu
Li, Yinchuan
Chen, Yingcong
contents We propose DiMeR, a novel geometry-texture disentangled feed-forward model with 3D supervision for sparse-view mesh reconstruction. Existing methods confront two persistent obstacles: (i) textures can conceal geometric errors, i.e., visually plausible images can be rendered even with wrong geometry, producing multiple ambiguous optimization objectives in geometry-texture mixed solution space for similar objects; and (ii) prevailing mesh extraction methods are redundant, unstable, and lack 3D supervision. To solve these challenges, we rethink the inductive bias for mesh reconstruction. First, we disentangle the unified geometry-texture solution space, where a single input admits multiple feasible solutions, into geometry and texture spaces individually. Specifically, given that normal maps are strictly consistent with geometry and accurately capture surface variations, the normal maps serve as the sole input for geometry prediction in DiMeR, while the texture is estimated from RGB images. Second, we streamline the algorithm of mesh extraction by eliminating modules with low performance/cost ratios and redesigning regularization losses with 3D supervision. Notably, DiMeR still accepts raw RGB images as input by leveraging foundation models for normal prediction. Extensive experiments demonstrate that DiMeR generalises across sparse-view-, single-image-, and text-to-3D tasks, consistently outperforming baselines. On the GSO and OmniObject3D datasets, DiMeR significantly reduces Chamfer Distance by more than 30%.
format Preprint
id arxiv_https___arxiv_org_abs_2504_17670
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DiMeR: Disentangled Mesh Reconstruction Model
Jiang, Lutao
Lin, Jiantao
Chen, Kanghao
Ge, Wenhang
Yang, Xin
Jiang, Yifan
Lyu, Yuanhuiyi
Zheng, Xu
Li, Yinchuan
Chen, Yingcong
Computer Vision and Pattern Recognition
We propose DiMeR, a novel geometry-texture disentangled feed-forward model with 3D supervision for sparse-view mesh reconstruction. Existing methods confront two persistent obstacles: (i) textures can conceal geometric errors, i.e., visually plausible images can be rendered even with wrong geometry, producing multiple ambiguous optimization objectives in geometry-texture mixed solution space for similar objects; and (ii) prevailing mesh extraction methods are redundant, unstable, and lack 3D supervision. To solve these challenges, we rethink the inductive bias for mesh reconstruction. First, we disentangle the unified geometry-texture solution space, where a single input admits multiple feasible solutions, into geometry and texture spaces individually. Specifically, given that normal maps are strictly consistent with geometry and accurately capture surface variations, the normal maps serve as the sole input for geometry prediction in DiMeR, while the texture is estimated from RGB images. Second, we streamline the algorithm of mesh extraction by eliminating modules with low performance/cost ratios and redesigning regularization losses with 3D supervision. Notably, DiMeR still accepts raw RGB images as input by leveraging foundation models for normal prediction. Extensive experiments demonstrate that DiMeR generalises across sparse-view-, single-image-, and text-to-3D tasks, consistently outperforming baselines. On the GSO and OmniObject3D datasets, DiMeR significantly reduces Chamfer Distance by more than 30%.
title DiMeR: Disentangled Mesh Reconstruction Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.17670