Del Visual al Auditivo: Sonorización de Escenas Guiada por Imagen

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Sánchez, María, Fernández, Laura, Arias, Julián, Cámara, Mateo, Comini, Giulia, Gabrys, Adam, Blanco, José Luis, Godino, Juan Ignacio, Hernández, Luis Alfonso
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916112784949248
author Sánchez, María
Fernández, Laura
Arias, Julián
Cámara, Mateo
Comini, Giulia
Gabrys, Adam
Blanco, José Luis
Godino, Juan Ignacio
Hernández, Luis Alfonso
author_facet Sánchez, María
Fernández, Laura
Arias, Julián
Cámara, Mateo
Comini, Giulia
Gabrys, Adam
Blanco, José Luis
Godino, Juan Ignacio
Hernández, Luis Alfonso
contents Recent advances in image, video, text and audio generative techniques, and their use by the general public, are leading to new forms of content generation. Usually, each modality was approached separately, which poses limitations. The automatic sound recording of visual sequences is one of the greatest challenges for the automatic generation of multimodal content. We present a processing flow that, starting from images extracted from videos, is able to sound them. We work with pre-trained models that employ complex encoders, contrastive learning, and multiple modalities, allowing complex representations of the sequences for their sonorization. The proposed scheme proposes different possibilities for audio mapping and text guidance. We evaluated the scheme on a dataset of frames extracted from a commercial video game and sounds extracted from the Freesound platform. Subjective tests have evidenced that the proposed scheme is able to generate and assign audios automatically and conveniently to images. Moreover, it adapts well to user preferences, and the proposed objective metrics show a high correlation with the subjective ratings.
format Preprint
id arxiv_https___arxiv_org_abs_2402_01385
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Del Visual al Auditivo: Sonorización de Escenas Guiada por Imagen
Sánchez, María
Fernández, Laura
Arias, Julián
Cámara, Mateo
Comini, Giulia
Gabrys, Adam
Blanco, José Luis
Godino, Juan Ignacio
Hernández, Luis Alfonso
Audio and Speech Processing
Sound
Recent advances in image, video, text and audio generative techniques, and their use by the general public, are leading to new forms of content generation. Usually, each modality was approached separately, which poses limitations. The automatic sound recording of visual sequences is one of the greatest challenges for the automatic generation of multimodal content. We present a processing flow that, starting from images extracted from videos, is able to sound them. We work with pre-trained models that employ complex encoders, contrastive learning, and multiple modalities, allowing complex representations of the sequences for their sonorization. The proposed scheme proposes different possibilities for audio mapping and text guidance. We evaluated the scheme on a dataset of frames extracted from a commercial video game and sounds extracted from the Freesound platform. Subjective tests have evidenced that the proposed scheme is able to generate and assign audios automatically and conveniently to images. Moreover, it adapts well to user preferences, and the proposed objective metrics show a high correlation with the subjective ratings.
title Del Visual al Auditivo: Sonorización de Escenas Guiada por Imagen
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2402.01385