Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chen, Jiahao, Yuan, Hangjie, Qian, Yichen, Liang, Jingyun, Xing, Jiazheng, Liu, Pengwei, Chen, Weihua, Wang, Fan, Su, Bing
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2506.02497
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916775090716672
author Chen, Jiahao
Yuan, Hangjie
Qian, Yichen
Liang, Jingyun
Xing, Jiazheng
Liu, Pengwei
Chen, Weihua
Wang, Fan
Su, Bing
author_facet Chen, Jiahao
Yuan, Hangjie
Qian, Yichen
Liang, Jingyun
Xing, Jiazheng
Liu, Pengwei
Chen, Weihua
Wang, Fan
Su, Bing
contents Long video generation has gained increasing attention due to its widespread applications in fields such as entertainment and simulation. Despite advances, synthesizing temporally coherent and visually compelling long sequences remains a formidable challenge. Conventional approaches often synthesize long videos by sequentially generating and concatenating short clips, or generating key frames and then interpolate the intermediate frames in a hierarchical manner. However, both of them still remain significant challenges, leading to issues such as temporal repetition or unnatural transitions. In this paper, we revisit the hierarchical long video generation pipeline and introduce LumosFlow, a framework introduce motion guidance explicitly. Specifically, we first employ the Large Motion Text-to-Video Diffusion Model (LMTV-DM) to generate key frames with larger motion intervals, thereby ensuring content diversity in the generated long videos. Given the complexity of interpolating contextual transitions between key frames, we further decompose the intermediate frame interpolation into motion generation and post-hoc refinement. For each pair of key frames, the Latent Optical Flow Diffusion Model (LOF-DM) synthesizes complex and large-motion optical flows, while MotionControlNet subsequently refines the warped results to enhance quality and guide intermediate frame generation. Compared with traditional video frame interpolation, we achieve 15x interpolation, ensuring reasonable and continuous motion between adjacent frames. Experiments show that our method can generate long videos with consistent motion and appearance. Code and models will be made publicly available upon acceptance. Our project page: https://jiahaochen1.github.io/LumosFlow/
format Preprint
id arxiv_https___arxiv_org_abs_2506_02497
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LumosFlow: Motion-Guided Long Video Generation
Chen, Jiahao
Yuan, Hangjie
Qian, Yichen
Liang, Jingyun
Xing, Jiazheng
Liu, Pengwei
Chen, Weihua
Wang, Fan
Su, Bing
Computer Vision and Pattern Recognition
Long video generation has gained increasing attention due to its widespread applications in fields such as entertainment and simulation. Despite advances, synthesizing temporally coherent and visually compelling long sequences remains a formidable challenge. Conventional approaches often synthesize long videos by sequentially generating and concatenating short clips, or generating key frames and then interpolate the intermediate frames in a hierarchical manner. However, both of them still remain significant challenges, leading to issues such as temporal repetition or unnatural transitions. In this paper, we revisit the hierarchical long video generation pipeline and introduce LumosFlow, a framework introduce motion guidance explicitly. Specifically, we first employ the Large Motion Text-to-Video Diffusion Model (LMTV-DM) to generate key frames with larger motion intervals, thereby ensuring content diversity in the generated long videos. Given the complexity of interpolating contextual transitions between key frames, we further decompose the intermediate frame interpolation into motion generation and post-hoc refinement. For each pair of key frames, the Latent Optical Flow Diffusion Model (LOF-DM) synthesizes complex and large-motion optical flows, while MotionControlNet subsequently refines the warped results to enhance quality and guide intermediate frame generation. Compared with traditional video frame interpolation, we achieve 15x interpolation, ensuring reasonable and continuous motion between adjacent frames. Experiments show that our method can generate long videos with consistent motion and appearance. Code and models will be made publicly available upon acceptance. Our project page: https://jiahaochen1.github.io/LumosFlow/
title LumosFlow: Motion-Guided Long Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.02497