Contrastive Sequential-Diffusion Learning: Non-linear and Multi-Scene Instructional Video Synthesis

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ramos, Vasco, Bitton, Yonatan, Yarom, Michal, Szpektor, Idan, Magalhaes, Joao
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909418857168896
author Ramos, Vasco
Bitton, Yonatan
Yarom, Michal
Szpektor, Idan
Magalhaes, Joao
author_facet Ramos, Vasco
Bitton, Yonatan
Yarom, Michal
Szpektor, Idan
Magalhaes, Joao
contents Generated video scenes for action-centric sequence descriptions, such as recipe instructions and do-it-yourself projects, often include non-linear patterns, where the next video may need to be visually consistent not with the immediately preceding video but with earlier ones. Current multi-scene video synthesis approaches fail to meet these consistency requirements. To address this, we propose a contrastive sequential video diffusion method that selects the most suitable previously generated scene to guide and condition the denoising process of the next scene. The result is a multi-scene video that is grounded in the scene descriptions and coherent w.r.t. the scenes that require visual consistency. Experiments with action-centered data from the real world demonstrate the practicality and improved consistency of our model compared to previous work.
format Preprint
id arxiv_https___arxiv_org_abs_2407_11814
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Contrastive Sequential-Diffusion Learning: Non-linear and Multi-Scene Instructional Video Synthesis
Ramos, Vasco
Bitton, Yonatan
Yarom, Michal
Szpektor, Idan
Magalhaes, Joao
Computer Vision and Pattern Recognition
Generated video scenes for action-centric sequence descriptions, such as recipe instructions and do-it-yourself projects, often include non-linear patterns, where the next video may need to be visually consistent not with the immediately preceding video but with earlier ones. Current multi-scene video synthesis approaches fail to meet these consistency requirements. To address this, we propose a contrastive sequential video diffusion method that selects the most suitable previously generated scene to guide and condition the denoising process of the next scene. The result is a multi-scene video that is grounded in the scene descriptions and coherent w.r.t. the scenes that require visual consistency. Experiments with action-centered data from the real world demonstrate the practicality and improved consistency of our model compared to previous work.
title Contrastive Sequential-Diffusion Learning: Non-linear and Multi-Scene Instructional Video Synthesis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.11814