Listen and Move: Improving GANs Coherency in Agnostic Sound-to-Video Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autore principale: Redondo, Rafael
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911930415841280
author Redondo, Rafael
author_facet Redondo, Rafael
contents Deep generative models have demonstrated the ability to create realistic audiovisual content, sometimes driven by domains of different nature. However, smooth temporal dynamics in video generation is a challenging problem. This work focuses on generic sound-to-video generation and proposes three main features to enhance both image quality and temporal coherency in generative adversarial models: a triple sound routing scheme, a multi-scale residual and dilated recurrent network for extended sound analysis, and a novel recurrent and directional convolutional layer for video prediction. Each of the proposed features improves, in both quality and coherency, the baseline neural architecture typically used in the SoTA, with the video prediction layer providing an extra temporal refinement.
format Preprint
id arxiv_https___arxiv_org_abs_2406_16155
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Listen and Move: Improving GANs Coherency in Agnostic Sound-to-Video Generation
Redondo, Rafael
Sound
Graphics
Audio and Speech Processing
Deep generative models have demonstrated the ability to create realistic audiovisual content, sometimes driven by domains of different nature. However, smooth temporal dynamics in video generation is a challenging problem. This work focuses on generic sound-to-video generation and proposes three main features to enhance both image quality and temporal coherency in generative adversarial models: a triple sound routing scheme, a multi-scale residual and dilated recurrent network for extended sound analysis, and a novel recurrent and directional convolutional layer for video prediction. Each of the proposed features improves, in both quality and coherency, the baseline neural architecture typically used in the SoTA, with the video prediction layer providing an extra temporal refinement.
title Listen and Move: Improving GANs Coherency in Agnostic Sound-to-Video Generation
topic Sound
Graphics
Audio and Speech Processing
url https://arxiv.org/abs/2406.16155