Contrastive Sequential-Diffusion Learning: Non-linear and Multi-Scene Instructional Video Synthesis
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909418857168896 |
|---|---|
| author | Ramos, Vasco Bitton, Yonatan Yarom, Michal Szpektor, Idan Magalhaes, Joao |
| author_facet | Ramos, Vasco Bitton, Yonatan Yarom, Michal Szpektor, Idan Magalhaes, Joao |
| contents | Generated video scenes for action-centric sequence descriptions, such as recipe instructions and do-it-yourself projects, often include non-linear patterns, where the next video may need to be visually consistent not with the immediately preceding video but with earlier ones. Current multi-scene video synthesis approaches fail to meet these consistency requirements. To address this, we propose a contrastive sequential video diffusion method that selects the most suitable previously generated scene to guide and condition the denoising process of the next scene. The result is a multi-scene video that is grounded in the scene descriptions and coherent w.r.t. the scenes that require visual consistency. Experiments with action-centered data from the real world demonstrate the practicality and improved consistency of our model compared to previous work. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_11814 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Contrastive Sequential-Diffusion Learning: Non-linear and Multi-Scene Instructional Video Synthesis Ramos, Vasco Bitton, Yonatan Yarom, Michal Szpektor, Idan Magalhaes, Joao Computer Vision and Pattern Recognition Generated video scenes for action-centric sequence descriptions, such as recipe instructions and do-it-yourself projects, often include non-linear patterns, where the next video may need to be visually consistent not with the immediately preceding video but with earlier ones. Current multi-scene video synthesis approaches fail to meet these consistency requirements. To address this, we propose a contrastive sequential video diffusion method that selects the most suitable previously generated scene to guide and condition the denoising process of the next scene. The result is a multi-scene video that is grounded in the scene descriptions and coherent w.r.t. the scenes that require visual consistency. Experiments with action-centered data from the real world demonstrate the practicality and improved consistency of our model compared to previous work. |
| title | Contrastive Sequential-Diffusion Learning: Non-linear and Multi-Scene Instructional Video Synthesis |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2407.11814 |