Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lokegaonkar, Vaibhavi, Bhosale, Aryan Vijay, Raj, Vishnu, KV, Gouthaman, Duraiswami, Ramani, Lu, Lie, Ghosh, Sreyan, Manocha, Dinesh
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910158330789888
author Lokegaonkar, Vaibhavi
Bhosale, Aryan Vijay
Raj, Vishnu
KV, Gouthaman
Duraiswami, Ramani
Lu, Lie
Ghosh, Sreyan
Manocha, Dinesh
author_facet Lokegaonkar, Vaibhavi
Bhosale, Aryan Vijay
Raj, Vishnu
KV, Gouthaman
Duraiswami, Ramani
Lu, Lie
Ghosh, Sreyan
Manocha, Dinesh
contents Video-to-music (V2M) is the fundamental task of creating background music for an input video. Recent V2M models achieve audiovisual alignment by typically relying on visual conditioning alone and provide limited semantic and stylistic controllability to the end user. In this paper, we present Video-Robin, a novel text-conditioned video-to-music generation model that enables fast, high-quality, semantically aligned music generation for video content. To balance musical fidelity and semantic understanding, Video-Robin integrates autoregressive planning with diffusion-based synthesis. Specifically, an autoregressive module models global structure by semantically aligning visual and textual inputs to produce high-level music latents. These latents are subsequently refined into coherent, high-fidelity music using local Diffusion Transformers. By factoring semantically driven planning into diffusion-based synthesis, Video-Robin enables fine-grained creator control without sacrificing audio realism. Our proposed model outperforms baselines that solely accept video input and additional feature conditioned baselines on both in-distribution and out-of-distribution benchmarks with a 2.21x speed in inference compared to SOTA. We will open-source everything upon paper acceptance.
format Preprint
id arxiv_https___arxiv_org_abs_2604_17656
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation
Lokegaonkar, Vaibhavi
Bhosale, Aryan Vijay
Raj, Vishnu
KV, Gouthaman
Duraiswami, Ramani
Lu, Lie
Ghosh, Sreyan
Manocha, Dinesh
Sound
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
Video-to-music (V2M) is the fundamental task of creating background music for an input video. Recent V2M models achieve audiovisual alignment by typically relying on visual conditioning alone and provide limited semantic and stylistic controllability to the end user. In this paper, we present Video-Robin, a novel text-conditioned video-to-music generation model that enables fast, high-quality, semantically aligned music generation for video content. To balance musical fidelity and semantic understanding, Video-Robin integrates autoregressive planning with diffusion-based synthesis. Specifically, an autoregressive module models global structure by semantically aligning visual and textual inputs to produce high-level music latents. These latents are subsequently refined into coherent, high-fidelity music using local Diffusion Transformers. By factoring semantically driven planning into diffusion-based synthesis, Video-Robin enables fine-grained creator control without sacrificing audio realism. Our proposed model outperforms baselines that solely accept video input and additional feature conditioned baselines on both in-distribution and out-of-distribution benchmarks with a 2.21x speed in inference compared to SOTA. We will open-source everything upon paper acceptance.
title Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation
topic Sound
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2604.17656