FoleyGRAM: Video-to-Audio Generation with GRAM-Aligned Multimodal Encoders
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866918155667898368 |
|---|---|
| author | Gramaccioni, Riccardo Fosco Marinoni, Christian Grassucci, Eleonora Cicchetti, Giordano Uncini, Aurelio Comminiello, Danilo |
| author_facet | Gramaccioni, Riccardo Fosco Marinoni, Christian Grassucci, Eleonora Cicchetti, Giordano Uncini, Aurelio Comminiello, Danilo |
| contents | In this work, we present FoleyGRAM, a novel approach to video-to-audio generation that emphasizes semantic conditioning through the use of aligned multimodal encoders. Building on prior advancements in video-to-audio generation, FoleyGRAM leverages the Gramian Representation Alignment Measure (GRAM) to align embeddings across video, text, and audio modalities, enabling precise semantic control over the audio generation process. The core of FoleyGRAM is a diffusion-based audio synthesis model conditioned on GRAM-aligned embeddings and waveform envelopes, ensuring both semantic richness and temporal alignment with the corresponding input video. We evaluate FoleyGRAM on the Greatest Hits dataset, a standard benchmark for video-to-audio models. Our experiments demonstrate that aligning multimodal encoders using GRAM enhances the system's ability to semantically align generated audio with video content, advancing the state of the art in video-to-audio synthesis. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_05829 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | FoleyGRAM: Video-to-Audio Generation with GRAM-Aligned Multimodal Encoders Gramaccioni, Riccardo Fosco Marinoni, Christian Grassucci, Eleonora Cicchetti, Giordano Uncini, Aurelio Comminiello, Danilo Sound Computer Vision and Pattern Recognition Machine Learning Multimedia Audio and Speech Processing In this work, we present FoleyGRAM, a novel approach to video-to-audio generation that emphasizes semantic conditioning through the use of aligned multimodal encoders. Building on prior advancements in video-to-audio generation, FoleyGRAM leverages the Gramian Representation Alignment Measure (GRAM) to align embeddings across video, text, and audio modalities, enabling precise semantic control over the audio generation process. The core of FoleyGRAM is a diffusion-based audio synthesis model conditioned on GRAM-aligned embeddings and waveform envelopes, ensuring both semantic richness and temporal alignment with the corresponding input video. We evaluate FoleyGRAM on the Greatest Hits dataset, a standard benchmark for video-to-audio models. Our experiments demonstrate that aligning multimodal encoders using GRAM enhances the system's ability to semantically align generated audio with video content, advancing the state of the art in video-to-audio synthesis. |
| title | FoleyGRAM: Video-to-Audio Generation with GRAM-Aligned Multimodal Encoders |
| topic | Sound Computer Vision and Pattern Recognition Machine Learning Multimedia Audio and Speech Processing |
| url | https://arxiv.org/abs/2510.05829 |