VidTune: Creating Video Soundtracks with Generative Music and Contextual Thumbnails

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huh, Mina, Fraser, C. Ailie, Li, Dingzeyu, Dontcheva, Mira, Wang, Bryan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911420339191808
author Huh, Mina
Fraser, C. Ailie
Li, Dingzeyu
Dontcheva, Mira
Wang, Bryan
author_facet Huh, Mina
Fraser, C. Ailie
Li, Dingzeyu
Dontcheva, Mira
Wang, Bryan
contents Music shapes the tone of videos, yet creators often struggle to find soundtracks that match their video's mood and narrative. Recent text-to-music models let creators generate music from text prompts, but our formative study (N=8) shows creators struggle to construct diverse prompts, quickly review and compare tracks, and understand their impact on the video. We present VidTune, a system that supports soundtrack creation by generating diverse music options from a creator's prompt and producing contextual thumbnails for rapid review. VidTune extracts representative video subjects to ground thumbnails in context, maps each track's valence and energy onto visual cues like color and brightness, and depicts prominent genres and instruments. Creators can refine tracks through natural language edits, which VidTune expands into new generations. In a controlled user study (N=12) and an exploratory case study (N=6), participants found VidTune helpful for efficiently reviewing and comparing music options and described the process as playful and enriching.
format Preprint
id arxiv_https___arxiv_org_abs_2601_12180
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VidTune: Creating Video Soundtracks with Generative Music and Contextual Thumbnails
Huh, Mina
Fraser, C. Ailie
Li, Dingzeyu
Dontcheva, Mira
Wang, Bryan
Human-Computer Interaction
Multimedia
Sound
Audio and Speech Processing
Music shapes the tone of videos, yet creators often struggle to find soundtracks that match their video's mood and narrative. Recent text-to-music models let creators generate music from text prompts, but our formative study (N=8) shows creators struggle to construct diverse prompts, quickly review and compare tracks, and understand their impact on the video. We present VidTune, a system that supports soundtrack creation by generating diverse music options from a creator's prompt and producing contextual thumbnails for rapid review. VidTune extracts representative video subjects to ground thumbnails in context, maps each track's valence and energy onto visual cues like color and brightness, and depicts prominent genres and instruments. Creators can refine tracks through natural language edits, which VidTune expands into new generations. In a controlled user study (N=12) and an exploratory case study (N=6), participants found VidTune helpful for efficiently reviewing and comparing music options and described the process as playful and enriching.
title VidTune: Creating Video Soundtracks with Generative Music and Contextual Thumbnails
topic Human-Computer Interaction
Multimedia
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2601.12180