When and Where do Events Switch in Multi-Event Video Generation?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liao, Ruotong, Huang, Guowen, Cheng, Qing, Seidl, Thomas, Cremers, Daniel, Tresp, Volker
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914072649269248
author Liao, Ruotong
Huang, Guowen
Cheng, Qing
Seidl, Thomas
Cremers, Daniel
Tresp, Volker
author_facet Liao, Ruotong
Huang, Guowen
Cheng, Qing
Seidl, Thomas
Cremers, Daniel
Tresp, Volker
contents Text-to-video (T2V) generation has surged in response to challenging questions, especially when a long video must depict multiple sequential events with temporal coherence and controllable content. Existing methods that extend to multi-event generation omit an inspection of the intrinsic factor in event shifting. The paper aims to answer the central question: When and where multi-event prompts control event transition during T2V generation. This work introduces MEve, a self-curated prompt suite for evaluating multi-event text-to-video (T2V) generation, and conducts a systematic study of two representative model families, i.e., OpenSora and CogVideoX. Extensive experiments demonstrate the importance of early intervention in denoising steps and block-wise model layers, revealing the essential factor for multi-event video generation and highlighting the possibilities for multi-event conditioning in future models.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03049
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When and Where do Events Switch in Multi-Event Video Generation?
Liao, Ruotong
Huang, Guowen
Cheng, Qing
Seidl, Thomas
Cremers, Daniel
Tresp, Volker
Computer Vision and Pattern Recognition
Artificial Intelligence
Text-to-video (T2V) generation has surged in response to challenging questions, especially when a long video must depict multiple sequential events with temporal coherence and controllable content. Existing methods that extend to multi-event generation omit an inspection of the intrinsic factor in event shifting. The paper aims to answer the central question: When and where multi-event prompts control event transition during T2V generation. This work introduces MEve, a self-curated prompt suite for evaluating multi-event text-to-video (T2V) generation, and conducts a systematic study of two representative model families, i.e., OpenSora and CogVideoX. Extensive experiments demonstrate the importance of early intervention in denoising steps and block-wise model layers, revealing the essential factor for multi-event video generation and highlighting the possibilities for multi-event conditioning in future models.
title When and Where do Events Switch in Multi-Event Video Generation?
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2510.03049