FoleyGRAM: Video-to-Audio Generation with GRAM-Aligned Multimodal Encoders

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gramaccioni, Riccardo Fosco, Marinoni, Christian, Grassucci, Eleonora, Cicchetti, Giordano, Uncini, Aurelio, Comminiello, Danilo
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918155667898368
author Gramaccioni, Riccardo Fosco
Marinoni, Christian
Grassucci, Eleonora
Cicchetti, Giordano
Uncini, Aurelio
Comminiello, Danilo
author_facet Gramaccioni, Riccardo Fosco
Marinoni, Christian
Grassucci, Eleonora
Cicchetti, Giordano
Uncini, Aurelio
Comminiello, Danilo
contents In this work, we present FoleyGRAM, a novel approach to video-to-audio generation that emphasizes semantic conditioning through the use of aligned multimodal encoders. Building on prior advancements in video-to-audio generation, FoleyGRAM leverages the Gramian Representation Alignment Measure (GRAM) to align embeddings across video, text, and audio modalities, enabling precise semantic control over the audio generation process. The core of FoleyGRAM is a diffusion-based audio synthesis model conditioned on GRAM-aligned embeddings and waveform envelopes, ensuring both semantic richness and temporal alignment with the corresponding input video. We evaluate FoleyGRAM on the Greatest Hits dataset, a standard benchmark for video-to-audio models. Our experiments demonstrate that aligning multimodal encoders using GRAM enhances the system's ability to semantically align generated audio with video content, advancing the state of the art in video-to-audio synthesis.
format Preprint
id arxiv_https___arxiv_org_abs_2510_05829
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FoleyGRAM: Video-to-Audio Generation with GRAM-Aligned Multimodal Encoders
Gramaccioni, Riccardo Fosco
Marinoni, Christian
Grassucci, Eleonora
Cicchetti, Giordano
Uncini, Aurelio
Comminiello, Danilo
Sound
Computer Vision and Pattern Recognition
Machine Learning
Multimedia
Audio and Speech Processing
In this work, we present FoleyGRAM, a novel approach to video-to-audio generation that emphasizes semantic conditioning through the use of aligned multimodal encoders. Building on prior advancements in video-to-audio generation, FoleyGRAM leverages the Gramian Representation Alignment Measure (GRAM) to align embeddings across video, text, and audio modalities, enabling precise semantic control over the audio generation process. The core of FoleyGRAM is a diffusion-based audio synthesis model conditioned on GRAM-aligned embeddings and waveform envelopes, ensuring both semantic richness and temporal alignment with the corresponding input video. We evaluate FoleyGRAM on the Greatest Hits dataset, a standard benchmark for video-to-audio models. Our experiments demonstrate that aligning multimodal encoders using GRAM enhances the system's ability to semantically align generated audio with video content, advancing the state of the art in video-to-audio synthesis.
title FoleyGRAM: Video-to-Audio Generation with GRAM-Aligned Multimodal Encoders
topic Sound
Computer Vision and Pattern Recognition
Machine Learning
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2510.05829