Saved in:
Bibliographic Details
Main Authors: Xiao, Junfei, Yang, Ceyuan, Zhang, Lvmin, Cai, Shengqu, Zhao, Yang, Guo, Yuwei, Wetzstein, Gordon, Agrawala, Maneesh, Yuille, Alan, Jiang, Lu
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2507.18634
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913959089537024
author Xiao, Junfei
Yang, Ceyuan
Zhang, Lvmin
Cai, Shengqu
Zhao, Yang
Guo, Yuwei
Wetzstein, Gordon
Agrawala, Maneesh
Yuille, Alan
Jiang, Lu
author_facet Xiao, Junfei
Yang, Ceyuan
Zhang, Lvmin
Cai, Shengqu
Zhao, Yang
Guo, Yuwei
Wetzstein, Gordon
Agrawala, Maneesh
Yuille, Alan
Jiang, Lu
contents We present Captain Cinema, a generation framework for short movie generation. Given a detailed textual description of a movie storyline, our approach firstly generates a sequence of keyframes that outline the entire narrative, which ensures long-range coherence in both the storyline and visual appearance (e.g., scenes and characters). We refer to this step as top-down keyframe planning. These keyframes then serve as conditioning signals for a video synthesis model, which supports long context learning, to produce the spatio-temporal dynamics between them. This step is referred to as bottom-up video synthesis. To support stable and efficient generation of multi-scene long narrative cinematic works, we introduce an interleaved training strategy for Multimodal Diffusion Transformers (MM-DiT), specifically adapted for long-context video data. Our model is trained on a specially curated cinematic dataset consisting of interleaved data pairs. Our experiments demonstrate that Captain Cinema performs favorably in the automated creation of visually coherent and narrative consistent short movies in high quality and efficiency. Project page: https://thecinema.ai
format Preprint
id arxiv_https___arxiv_org_abs_2507_18634
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Captain Cinema: Towards Short Movie Generation
Xiao, Junfei
Yang, Ceyuan
Zhang, Lvmin
Cai, Shengqu
Zhao, Yang
Guo, Yuwei
Wetzstein, Gordon
Agrawala, Maneesh
Yuille, Alan
Jiang, Lu
Computer Vision and Pattern Recognition
We present Captain Cinema, a generation framework for short movie generation. Given a detailed textual description of a movie storyline, our approach firstly generates a sequence of keyframes that outline the entire narrative, which ensures long-range coherence in both the storyline and visual appearance (e.g., scenes and characters). We refer to this step as top-down keyframe planning. These keyframes then serve as conditioning signals for a video synthesis model, which supports long context learning, to produce the spatio-temporal dynamics between them. This step is referred to as bottom-up video synthesis. To support stable and efficient generation of multi-scene long narrative cinematic works, we introduce an interleaved training strategy for Multimodal Diffusion Transformers (MM-DiT), specifically adapted for long-context video data. Our model is trained on a specially curated cinematic dataset consisting of interleaved data pairs. Our experiments demonstrate that Captain Cinema performs favorably in the automated creation of visually coherent and narrative consistent short movies in high quality and efficiency. Project page: https://thecinema.ai
title Captain Cinema: Towards Short Movie Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.18634