Light-X: Generative 4D Video Rendering with Camera and Illumination Control

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Tianqi, Chen, Zhaoxi, Huang, Zihao, Xu, Shaocong, Zhang, Saining, Ye, Chongjie, Li, Bohan, Cao, Zhiguo, Li, Wei, Zhao, Hao, Liu, Ziwei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918248738455552
author Liu, Tianqi
Chen, Zhaoxi
Huang, Zihao
Xu, Shaocong
Zhang, Saining
Ye, Chongjie
Li, Bohan
Cao, Zhiguo
Li, Wei
Zhao, Hao
Liu, Ziwei
author_facet Liu, Tianqi
Chen, Zhaoxi
Huang, Zihao
Xu, Shaocong
Zhang, Saining
Ye, Chongjie
Li, Bohan
Cao, Zhiguo
Li, Wei
Zhao, Hao
Liu, Ziwei
contents Recent advances in illumination control extend image-based methods to video, yet still facing a trade-off between lighting fidelity and temporal consistency. Moving beyond relighting, a key step toward generative modeling of real-world scenes is the joint control of camera trajectory and illumination, since visual dynamics are inherently shaped by both geometry and lighting. To this end, we present Light-X, a video generation framework that enables controllable rendering from monocular videos with both viewpoint and illumination control. 1) We propose a disentangled design that decouples geometry and lighting signals: geometry and motion are captured via dynamic point clouds projected along user-defined camera trajectories, while illumination cues are provided by a relit frame consistently projected into the same geometry. These explicit, fine-grained cues enable effective disentanglement and guide high-quality illumination. 2) To address the lack of paired multi-view and multi-illumination videos, we introduce Light-Syn, a degradation-based pipeline with inverse-mapping that synthesizes training pairs from in-the-wild monocular footage. This strategy yields a dataset covering static, dynamic, and AI-generated scenes, ensuring robust training. Extensive experiments show that Light-X outperforms baseline methods in joint camera-illumination control and surpasses prior video relighting methods under both text- and background-conditioned settings.
format Preprint
id arxiv_https___arxiv_org_abs_2512_05115
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Light-X: Generative 4D Video Rendering with Camera and Illumination Control
Liu, Tianqi
Chen, Zhaoxi
Huang, Zihao
Xu, Shaocong
Zhang, Saining
Ye, Chongjie
Li, Bohan
Cao, Zhiguo
Li, Wei
Zhao, Hao
Liu, Ziwei
Computer Vision and Pattern Recognition
Recent advances in illumination control extend image-based methods to video, yet still facing a trade-off between lighting fidelity and temporal consistency. Moving beyond relighting, a key step toward generative modeling of real-world scenes is the joint control of camera trajectory and illumination, since visual dynamics are inherently shaped by both geometry and lighting. To this end, we present Light-X, a video generation framework that enables controllable rendering from monocular videos with both viewpoint and illumination control. 1) We propose a disentangled design that decouples geometry and lighting signals: geometry and motion are captured via dynamic point clouds projected along user-defined camera trajectories, while illumination cues are provided by a relit frame consistently projected into the same geometry. These explicit, fine-grained cues enable effective disentanglement and guide high-quality illumination. 2) To address the lack of paired multi-view and multi-illumination videos, we introduce Light-Syn, a degradation-based pipeline with inverse-mapping that synthesizes training pairs from in-the-wild monocular footage. This strategy yields a dataset covering static, dynamic, and AI-generated scenes, ensuring robust training. Extensive experiments show that Light-X outperforms baseline methods in joint camera-illumination control and surpasses prior video relighting methods under both text- and background-conditioned settings.
title Light-X: Generative 4D Video Rendering with Camera and Illumination Control
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.05115