DreaMontage: Arbitrary Frame-Guided One-Shot Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Jiawei, Li, Junqiao, Deng, Jiangfan, Li, Gen, Zhou, Siyu, Fang, Zetao, Lao, Shanshan, Deng, Zengde, Zhu, Jianing, Ma, Tingting, Li, Jiayi, Wang, Yunqiu, He, Qian, Wu, Xinglong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909975731765248
author Liu, Jiawei
Li, Junqiao
Deng, Jiangfan
Li, Gen
Zhou, Siyu
Fang, Zetao
Lao, Shanshan
Deng, Zengde
Zhu, Jianing
Ma, Tingting
Li, Jiayi
Wang, Yunqiu
He, Qian
Wu, Xinglong
author_facet Liu, Jiawei
Li, Junqiao
Deng, Jiangfan
Li, Gen
Zhou, Siyu
Fang, Zetao
Lao, Shanshan
Deng, Zengde
Zhu, Jianing
Ma, Tingting
Li, Jiayi
Wang, Yunqiu
He, Qian
Wu, Xinglong
contents The "one-shot" technique represents a distinct and sophisticated aesthetic in filmmaking. However, its practical realization is often hindered by prohibitive costs and complex real-world constraints. Although emerging video generation models offer a virtual alternative, existing approaches typically rely on naive clip concatenation, which frequently fails to maintain visual smoothness and temporal coherence. In this paper, we introduce DreaMontage, a comprehensive framework designed for arbitrary frame-guided generation, capable of synthesizing seamless, expressive, and long-duration one-shot videos from diverse user-provided inputs. To achieve this, we address the challenge through three primary dimensions. (i) We integrate a lightweight intermediate-conditioning mechanism into the DiT architecture. By employing an Adaptive Tuning strategy that effectively leverages base training data, we unlock robust arbitrary-frame control capabilities. (ii) To enhance visual fidelity and cinematic expressiveness, we curate a high-quality dataset and implement a Visual Expression SFT stage. In addressing critical issues such as subject motion rationality and transition smoothness, we apply a Tailored DPO scheme, which significantly improves the success rate and usability of the generated content. (iii) To facilitate the production of extended sequences, we design a Segment-wise Auto-Regressive (SAR) inference strategy that operates in a memory-efficient manner. Extensive experiments demonstrate that our approach achieves visually striking and seamlessly coherent one-shot effects while maintaining computational efficiency, empowering users to transform fragmented visual materials into vivid, cohesive one-shot cinematic experiences.
format Preprint
id arxiv_https___arxiv_org_abs_2512_21252
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DreaMontage: Arbitrary Frame-Guided One-Shot Video Generation
Liu, Jiawei
Li, Junqiao
Deng, Jiangfan
Li, Gen
Zhou, Siyu
Fang, Zetao
Lao, Shanshan
Deng, Zengde
Zhu, Jianing
Ma, Tingting
Li, Jiayi
Wang, Yunqiu
He, Qian
Wu, Xinglong
Computer Vision and Pattern Recognition
The "one-shot" technique represents a distinct and sophisticated aesthetic in filmmaking. However, its practical realization is often hindered by prohibitive costs and complex real-world constraints. Although emerging video generation models offer a virtual alternative, existing approaches typically rely on naive clip concatenation, which frequently fails to maintain visual smoothness and temporal coherence. In this paper, we introduce DreaMontage, a comprehensive framework designed for arbitrary frame-guided generation, capable of synthesizing seamless, expressive, and long-duration one-shot videos from diverse user-provided inputs. To achieve this, we address the challenge through three primary dimensions. (i) We integrate a lightweight intermediate-conditioning mechanism into the DiT architecture. By employing an Adaptive Tuning strategy that effectively leverages base training data, we unlock robust arbitrary-frame control capabilities. (ii) To enhance visual fidelity and cinematic expressiveness, we curate a high-quality dataset and implement a Visual Expression SFT stage. In addressing critical issues such as subject motion rationality and transition smoothness, we apply a Tailored DPO scheme, which significantly improves the success rate and usability of the generated content. (iii) To facilitate the production of extended sequences, we design a Segment-wise Auto-Regressive (SAR) inference strategy that operates in a memory-efficient manner. Extensive experiments demonstrate that our approach achieves visually striking and seamlessly coherent one-shot effects while maintaining computational efficiency, empowering users to transform fragmented visual materials into vivid, cohesive one-shot cinematic experiences.
title DreaMontage: Arbitrary Frame-Guided One-Shot Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.21252