MVInverse: Feed-forward Multi-view Inverse Rendering in Seconds

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Xiangzuo, Ren, Chengwei, Zhou, Jun, Li, Xiu, Liu, Yuan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909977387466752
author Wu, Xiangzuo
Ren, Chengwei
Zhou, Jun
Li, Xiu
Liu, Yuan
author_facet Wu, Xiangzuo
Ren, Chengwei
Zhou, Jun
Li, Xiu
Liu, Yuan
contents Multi-view inverse rendering aims to recover geometry, materials, and illumination consistently across multiple viewpoints. When applied to multi-view images, existing single-view approaches often ignore cross-view relationships, leading to inconsistent results. In contrast, multi-view optimization methods rely on slow differentiable rendering and per-scene refinement, making them computationally expensive and hard to scale. To address these limitations, we introduce a feed-forward multi-view inverse rendering framework that directly predicts spatially varying albedo, metallic, roughness, diffuse shading, and surface normals from sequences of RGB images. By alternating attention across views, our model captures both intra-view long-range lighting interactions and inter-view material consistency, enabling coherent scene-level reasoning within a single forward pass. Due to the scarcity of real-world training data, models trained on existing synthetic datasets often struggle to generalize to real-world scenes. To overcome this limitation, we propose a consistency-based finetuning strategy that leverages unlabeled real-world videos to enhance both multi-view coherence and robustness under in-the-wild conditions. Extensive experiments on benchmark datasets demonstrate that our method achieves state-of-the-art performance in terms of multi-view consistency, material and normal estimation quality, and generalization to real-world imagery. Project page: https://maddog241.github.io/mvinverse-page/
format Preprint
id arxiv_https___arxiv_org_abs_2512_21003
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MVInverse: Feed-forward Multi-view Inverse Rendering in Seconds
Wu, Xiangzuo
Ren, Chengwei
Zhou, Jun
Li, Xiu
Liu, Yuan
Computer Vision and Pattern Recognition
Multi-view inverse rendering aims to recover geometry, materials, and illumination consistently across multiple viewpoints. When applied to multi-view images, existing single-view approaches often ignore cross-view relationships, leading to inconsistent results. In contrast, multi-view optimization methods rely on slow differentiable rendering and per-scene refinement, making them computationally expensive and hard to scale. To address these limitations, we introduce a feed-forward multi-view inverse rendering framework that directly predicts spatially varying albedo, metallic, roughness, diffuse shading, and surface normals from sequences of RGB images. By alternating attention across views, our model captures both intra-view long-range lighting interactions and inter-view material consistency, enabling coherent scene-level reasoning within a single forward pass. Due to the scarcity of real-world training data, models trained on existing synthetic datasets often struggle to generalize to real-world scenes. To overcome this limitation, we propose a consistency-based finetuning strategy that leverages unlabeled real-world videos to enhance both multi-view coherence and robustness under in-the-wild conditions. Extensive experiments on benchmark datasets demonstrate that our method achieves state-of-the-art performance in terms of multi-view consistency, material and normal estimation quality, and generalization to real-world imagery. Project page: https://maddog241.github.io/mvinverse-page/
title MVInverse: Feed-forward Multi-view Inverse Rendering in Seconds
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.21003