LongVie: Multimodal-Guided Controllable Ultra-Long Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Jianxiong, Chen, Zhaoxi, Liu, Xian, Feng, Jianfeng, Si, Chenyang, Fu, Yanwei, Qiao, Yu, Liu, Ziwei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915429056774144
author Gao, Jianxiong
Chen, Zhaoxi
Liu, Xian
Feng, Jianfeng
Si, Chenyang
Fu, Yanwei
Qiao, Yu
Liu, Ziwei
author_facet Gao, Jianxiong
Chen, Zhaoxi
Liu, Xian
Feng, Jianfeng
Si, Chenyang
Fu, Yanwei
Qiao, Yu
Liu, Ziwei
contents Controllable ultra-long video generation is a fundamental yet challenging task. Although existing methods are effective for short clips, they struggle to scale due to issues such as temporal inconsistency and visual degradation. In this paper, we initially investigate and identify three key factors: separate noise initialization, independent control signal normalization, and the limitations of single-modality guidance. To address these issues, we propose LongVie, an end-to-end autoregressive framework for controllable long video generation. LongVie introduces two core designs to ensure temporal consistency: 1) a unified noise initialization strategy that maintains consistent generation across clips, and 2) global control signal normalization that enforces alignment in the control space throughout the entire video. To mitigate visual degradation, LongVie employs 3) a multi-modal control framework that integrates both dense (e.g., depth maps) and sparse (e.g., keypoints) control signals, complemented by 4) a degradation-aware training strategy that adaptively balances modality contributions over time to preserve visual quality. We also introduce LongVGenBench, a comprehensive benchmark consisting of 100 high-resolution videos spanning diverse real-world and synthetic environments, each lasting over one minute. Extensive experiments show that LongVie achieves state-of-the-art performance in long-range controllability, consistency, and quality.
format Preprint
id arxiv_https___arxiv_org_abs_2508_03694
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LongVie: Multimodal-Guided Controllable Ultra-Long Video Generation
Gao, Jianxiong
Chen, Zhaoxi
Liu, Xian
Feng, Jianfeng
Si, Chenyang
Fu, Yanwei
Qiao, Yu
Liu, Ziwei
Computer Vision and Pattern Recognition
Controllable ultra-long video generation is a fundamental yet challenging task. Although existing methods are effective for short clips, they struggle to scale due to issues such as temporal inconsistency and visual degradation. In this paper, we initially investigate and identify three key factors: separate noise initialization, independent control signal normalization, and the limitations of single-modality guidance. To address these issues, we propose LongVie, an end-to-end autoregressive framework for controllable long video generation. LongVie introduces two core designs to ensure temporal consistency: 1) a unified noise initialization strategy that maintains consistent generation across clips, and 2) global control signal normalization that enforces alignment in the control space throughout the entire video. To mitigate visual degradation, LongVie employs 3) a multi-modal control framework that integrates both dense (e.g., depth maps) and sparse (e.g., keypoints) control signals, complemented by 4) a degradation-aware training strategy that adaptively balances modality contributions over time to preserve visual quality. We also introduce LongVGenBench, a comprehensive benchmark consisting of 100 high-resolution videos spanning diverse real-world and synthetic environments, each lasting over one minute. Extensive experiments show that LongVie achieves state-of-the-art performance in long-range controllability, consistency, and quality.
title LongVie: Multimodal-Guided Controllable Ultra-Long Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.03694