CamLit: Unified Video Diffusion with Explicit Camera and Lighting Control

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kuang, Zhiyi, He, Chengan, Zakharov, Egor, Xue, Yuxuan, Saito, Shunsuke, Maury, Olivier, Bagautdinov, Timur, Zheng, Youyi, Nam, Giljoo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918389426946048
author Kuang, Zhiyi
He, Chengan
Zakharov, Egor
Xue, Yuxuan
Saito, Shunsuke
Maury, Olivier
Bagautdinov, Timur
Zheng, Youyi
Nam, Giljoo
author_facet Kuang, Zhiyi
He, Chengan
Zakharov, Egor
Xue, Yuxuan
Saito, Shunsuke
Maury, Olivier
Bagautdinov, Timur
Zheng, Youyi
Nam, Giljoo
contents We present CamLit, the first unified video diffusion model that jointly performs novel view synthesis (NVS) and relighting from a single input image. Given one reference image, a user-defined camera trajectory, and an environment map, CamLit synthesizes a video of the scene from new viewpoints under the specified illumination. Within a single generative process, our model produces temporally coherent and spatially aligned outputs, including relit novel-view frames and corresponding albedo frames, enabling high-quality control of both camera pose and lighting. Qualitative and quantitative experiments demonstrate that CamLit achieves high-fidelity outputs on par with state-of-the-art methods in both novel view synthesis and relighting, without sacrificing visual quality in either task. We show that a single generative model can effectively integrate camera and lighting control, simplifying the video generation pipeline while maintaining competitive performance and consistent realism.
format Preprint
id arxiv_https___arxiv_org_abs_2603_14241
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CamLit: Unified Video Diffusion with Explicit Camera and Lighting Control
Kuang, Zhiyi
He, Chengan
Zakharov, Egor
Xue, Yuxuan
Saito, Shunsuke
Maury, Olivier
Bagautdinov, Timur
Zheng, Youyi
Nam, Giljoo
Computer Vision and Pattern Recognition
We present CamLit, the first unified video diffusion model that jointly performs novel view synthesis (NVS) and relighting from a single input image. Given one reference image, a user-defined camera trajectory, and an environment map, CamLit synthesizes a video of the scene from new viewpoints under the specified illumination. Within a single generative process, our model produces temporally coherent and spatially aligned outputs, including relit novel-view frames and corresponding albedo frames, enabling high-quality control of both camera pose and lighting. Qualitative and quantitative experiments demonstrate that CamLit achieves high-fidelity outputs on par with state-of-the-art methods in both novel view synthesis and relighting, without sacrificing visual quality in either task. We show that a single generative model can effectively integrate camera and lighting control, simplifying the video generation pipeline while maintaining competitive performance and consistent realism.
title CamLit: Unified Video Diffusion with Explicit Camera and Lighting Control
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.14241