GeoRelight: Learning Joint Geometrical Relighting and Reconstruction with Flexible Multi-Modal Diffusion Transformers

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xue, Yuxuan, Liang, Ruofan, Zakharov, Egor, Bagautdinov, Timur, Cao, Chen, Nam, Giljoo, Saito, Shunsuke, Pons-Moll, Gerard, Romero, Javier
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918462070194176
author Xue, Yuxuan
Liang, Ruofan
Zakharov, Egor
Bagautdinov, Timur
Cao, Chen
Nam, Giljoo
Saito, Shunsuke
Pons-Moll, Gerard
Romero, Javier
author_facet Xue, Yuxuan
Liang, Ruofan
Zakharov, Egor
Bagautdinov, Timur
Cao, Chen
Nam, Giljoo
Saito, Shunsuke
Pons-Moll, Gerard
Romero, Javier
contents Relighting a person from a single photo is an attractive but ill-posed task, as a 2D image ambiguously entangles 3D geometry, intrinsic appearance, and illumination. Current methods either use sequential pipelines that suffer from error accumulation, or they do not explicitly leverage 3D geometry during relighting, which limits physical consistency. Since relighting and estimation of 3D geometry are mutually beneficial tasks, we propose a unified Multi-Modal Diffusion Transformer (DiT) that jointly solves for both: GeoRelight. We make this possible through two key technical contributions: isotropic NDC-Orthographic Depth (iNOD), a distortion-free 3D representation compatible with latent diffusion models; and a strategic mixed-data training method that combines synthetic and auto-labeled real data. By solving geometry and relighting jointly, GeoRelight achieves better performance than both sequential models and previous systems that ignored geometry.
format Preprint
id arxiv_https___arxiv_org_abs_2604_20715
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GeoRelight: Learning Joint Geometrical Relighting and Reconstruction with Flexible Multi-Modal Diffusion Transformers
Xue, Yuxuan
Liang, Ruofan
Zakharov, Egor
Bagautdinov, Timur
Cao, Chen
Nam, Giljoo
Saito, Shunsuke
Pons-Moll, Gerard
Romero, Javier
Computer Vision and Pattern Recognition
Relighting a person from a single photo is an attractive but ill-posed task, as a 2D image ambiguously entangles 3D geometry, intrinsic appearance, and illumination. Current methods either use sequential pipelines that suffer from error accumulation, or they do not explicitly leverage 3D geometry during relighting, which limits physical consistency. Since relighting and estimation of 3D geometry are mutually beneficial tasks, we propose a unified Multi-Modal Diffusion Transformer (DiT) that jointly solves for both: GeoRelight. We make this possible through two key technical contributions: isotropic NDC-Orthographic Depth (iNOD), a distortion-free 3D representation compatible with latent diffusion models; and a strategic mixed-data training method that combines synthetic and auto-labeled real data. By solving geometry and relighting jointly, GeoRelight achieves better performance than both sequential models and previous systems that ignored geometry.
title GeoRelight: Learning Joint Geometrical Relighting and Reconstruction with Flexible Multi-Modal Diffusion Transformers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.20715