LGTM: Training-Free Light-Guided Text-to-Image Diffusion Model via Initial Noise Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Morita, Ryugo, Frolov, Stanislav, Moser, Brian Bernhard, Watanabe, Ko, Takahashi, Riku, Dengel, Andreas
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914420579368960
author Morita, Ryugo
Frolov, Stanislav
Moser, Brian Bernhard
Watanabe, Ko
Takahashi, Riku
Dengel, Andreas
author_facet Morita, Ryugo
Frolov, Stanislav
Moser, Brian Bernhard
Watanabe, Ko
Takahashi, Riku
Dengel, Andreas
contents Diffusion models have demonstrated high-quality performance in conditional text-to-image generation, particularly with structural cues such as edges, layouts, and depth. However, lighting conditions have received limited attention and remain difficult to control within the generative process. Existing methods handle lighting through a two-stage pipeline that relights images after generation, which is inefficient. Moreover, they rely on fine-tuning with large datasets and heavy computation, limiting their adaptability to new models and tasks. To address this, we propose a novel Training-Free Light-Guided Text-to-Image Diffusion Model via Initial Noise Manipulation (LGTM), which manipulates the initial latent noise of the diffusion process to guide image generation with text prompts and user-specified light directions. Through a channel-wise analysis of the latent space, we find that selectively manipulating latent channels enables fine-grained lighting control without fine-tuning or modifying the pre-trained model. Extensive experiments show that our method surpasses prompt-based baselines in lighting consistency, while preserving image quality and text alignment. This approach introduces new possibilities for dynamic, user-guided light control. Furthermore, it integrates seamlessly with models like ControlNet, demonstrating adaptability across diverse scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2603_24086
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LGTM: Training-Free Light-Guided Text-to-Image Diffusion Model via Initial Noise Manipulation
Morita, Ryugo
Frolov, Stanislav
Moser, Brian Bernhard
Watanabe, Ko
Takahashi, Riku
Dengel, Andreas
Computer Vision and Pattern Recognition
Graphics
Diffusion models have demonstrated high-quality performance in conditional text-to-image generation, particularly with structural cues such as edges, layouts, and depth. However, lighting conditions have received limited attention and remain difficult to control within the generative process. Existing methods handle lighting through a two-stage pipeline that relights images after generation, which is inefficient. Moreover, they rely on fine-tuning with large datasets and heavy computation, limiting their adaptability to new models and tasks. To address this, we propose a novel Training-Free Light-Guided Text-to-Image Diffusion Model via Initial Noise Manipulation (LGTM), which manipulates the initial latent noise of the diffusion process to guide image generation with text prompts and user-specified light directions. Through a channel-wise analysis of the latent space, we find that selectively manipulating latent channels enables fine-grained lighting control without fine-tuning or modifying the pre-trained model. Extensive experiments show that our method surpasses prompt-based baselines in lighting consistency, while preserving image quality and text alignment. This approach introduces new possibilities for dynamic, user-guided light control. Furthermore, it integrates seamlessly with models like ControlNet, demonstrating adaptability across diverse scenarios.
title LGTM: Training-Free Light-Guided Text-to-Image Diffusion Model via Initial Noise Manipulation
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2603.24086