LumiX: Structured and Coherent Text-to-Intrinsic Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Han, Xu, Zhang, Biao, Tang, Xiangjun, Li, Xianzhi, Wonka, Peter
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914177623261184
author Han, Xu
Zhang, Biao
Tang, Xiangjun
Li, Xianzhi
Wonka, Peter
author_facet Han, Xu
Zhang, Biao
Tang, Xiangjun
Li, Xianzhi
Wonka, Peter
contents We present LumiX, a structured diffusion framework for coherent text-to-intrinsic generation. Conditioned on text prompts, LumiX jointly generates a comprehensive set of intrinsic maps (e.g., albedo, irradiance, normal, depth, and final color), providing a structured and physically consistent description of an underlying scene. This is enabled by two key contributions: 1) Query-Broadcast Attention, a mechanism that ensures structural consistency by sharing queries across all maps in each self-attention block. 2) Tensor LoRA, a tensor-based adaptation that parameter-efficiently models cross-map relations for efficient joint training. Together, these designs enable stable joint diffusion training and unified generation of multiple intrinsic properties. Experiments show that LumiX produces coherent and physically meaningful results, achieving 23% higher alignment and a better preference score (0.19 vs. -0.41) compared to the state of the art, and it can also perform image-conditioned intrinsic decomposition within the same framework.
format Preprint
id arxiv_https___arxiv_org_abs_2512_02781
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LumiX: Structured and Coherent Text-to-Intrinsic Generation
Han, Xu
Zhang, Biao
Tang, Xiangjun
Li, Xianzhi
Wonka, Peter
Computer Vision and Pattern Recognition
Graphics
Machine Learning
We present LumiX, a structured diffusion framework for coherent text-to-intrinsic generation. Conditioned on text prompts, LumiX jointly generates a comprehensive set of intrinsic maps (e.g., albedo, irradiance, normal, depth, and final color), providing a structured and physically consistent description of an underlying scene. This is enabled by two key contributions: 1) Query-Broadcast Attention, a mechanism that ensures structural consistency by sharing queries across all maps in each self-attention block. 2) Tensor LoRA, a tensor-based adaptation that parameter-efficiently models cross-map relations for efficient joint training. Together, these designs enable stable joint diffusion training and unified generation of multiple intrinsic properties. Experiments show that LumiX produces coherent and physically meaningful results, achieving 23% higher alignment and a better preference score (0.19 vs. -0.41) compared to the state of the art, and it can also perform image-conditioned intrinsic decomposition within the same framework.
title LumiX: Structured and Coherent Text-to-Intrinsic Generation
topic Computer Vision and Pattern Recognition
Graphics
Machine Learning
url https://arxiv.org/abs/2512.02781