CineVerse: Consistent Keyframe Synthesis for Cinematic Scene Composition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Phung, Quynh, Mai, Long, Heilbron, Fabian David Caba, Liu, Feng, Huang, Jia-Bin, Ham, Cusuh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912350515232768
author Phung, Quynh
Mai, Long
Heilbron, Fabian David Caba
Liu, Feng
Huang, Jia-Bin
Ham, Cusuh
author_facet Phung, Quynh
Mai, Long
Heilbron, Fabian David Caba
Liu, Feng
Huang, Jia-Bin
Ham, Cusuh
contents We present CineVerse, a novel framework for the task of cinematic scene composition. Similar to traditional multi-shot generation, our task emphasizes the need for consistency and continuity across frames. However, our task also focuses on addressing challenges inherent to filmmaking, such as multiple characters, complex interactions, and visual cinematic effects. In order to learn to generate such content, we first create the CineVerse dataset. We use this dataset to train our proposed two-stage approach. First, we prompt a large language model (LLM) with task-specific instructions to take in a high-level scene description and generate a detailed plan for the overall setting and characters, as well as the individual shots. Then, we fine-tune a text-to-image generation model to synthesize high-quality visual keyframes. Experimental results demonstrate that CineVerse yields promising improvements in generating visually coherent and contextually rich movie scenes, paving the way for further exploration in cinematic video synthesis.
format Preprint
id arxiv_https___arxiv_org_abs_2504_19894
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CineVerse: Consistent Keyframe Synthesis for Cinematic Scene Composition
Phung, Quynh
Mai, Long
Heilbron, Fabian David Caba
Liu, Feng
Huang, Jia-Bin
Ham, Cusuh
Computer Vision and Pattern Recognition
We present CineVerse, a novel framework for the task of cinematic scene composition. Similar to traditional multi-shot generation, our task emphasizes the need for consistency and continuity across frames. However, our task also focuses on addressing challenges inherent to filmmaking, such as multiple characters, complex interactions, and visual cinematic effects. In order to learn to generate such content, we first create the CineVerse dataset. We use this dataset to train our proposed two-stage approach. First, we prompt a large language model (LLM) with task-specific instructions to take in a high-level scene description and generate a detailed plan for the overall setting and characters, as well as the individual shots. Then, we fine-tune a text-to-image generation model to synthesize high-quality visual keyframes. Experimental results demonstrate that CineVerse yields promising improvements in generating visually coherent and contextually rich movie scenes, paving the way for further exploration in cinematic video synthesis.
title CineVerse: Consistent Keyframe Synthesis for Cinematic Scene Composition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.19894