When and Where do Events Switch in Multi-Event Video Generation?
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914072649269248 |
|---|---|
| author | Liao, Ruotong Huang, Guowen Cheng, Qing Seidl, Thomas Cremers, Daniel Tresp, Volker |
| author_facet | Liao, Ruotong Huang, Guowen Cheng, Qing Seidl, Thomas Cremers, Daniel Tresp, Volker |
| contents | Text-to-video (T2V) generation has surged in response to challenging questions, especially when a long video must depict multiple sequential events with temporal coherence and controllable content. Existing methods that extend to multi-event generation omit an inspection of the intrinsic factor in event shifting. The paper aims to answer the central question: When and where multi-event prompts control event transition during T2V generation. This work introduces MEve, a self-curated prompt suite for evaluating multi-event text-to-video (T2V) generation, and conducts a systematic study of two representative model families, i.e., OpenSora and CogVideoX. Extensive experiments demonstrate the importance of early intervention in denoising steps and block-wise model layers, revealing the essential factor for multi-event video generation and highlighting the possibilities for multi-event conditioning in future models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_03049 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | When and Where do Events Switch in Multi-Event Video Generation? Liao, Ruotong Huang, Guowen Cheng, Qing Seidl, Thomas Cremers, Daniel Tresp, Volker Computer Vision and Pattern Recognition Artificial Intelligence Text-to-video (T2V) generation has surged in response to challenging questions, especially when a long video must depict multiple sequential events with temporal coherence and controllable content. Existing methods that extend to multi-event generation omit an inspection of the intrinsic factor in event shifting. The paper aims to answer the central question: When and where multi-event prompts control event transition during T2V generation. This work introduces MEve, a self-curated prompt suite for evaluating multi-event text-to-video (T2V) generation, and conducts a systematic study of two representative model families, i.e., OpenSora and CogVideoX. Extensive experiments demonstrate the importance of early intervention in denoising steps and block-wise model layers, revealing the essential factor for multi-event video generation and highlighting the possibilities for multi-event conditioning in future models. |
| title | When and Where do Events Switch in Multi-Event Video Generation? |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2510.03049 |