Multi Agents Semantic Emotion Aligned Music to Image Generation with Music Derived Captions

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Shi, Junchang, Li, Gang
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909977623396352
author Shi, Junchang
Li, Gang
author_facet Shi, Junchang
Li, Gang
contents When people listen to music, they often experience rich visual imagery. We aim to externalize this inner imagery by generating images conditioned on music. We propose MESA MIG, a multi agent semantic and emotion aligned framework that first produces structured music captions and then refines them with cooperating agents specializing in scene, motion, style, color, and composition. In parallel, a Valence Arousal regression head predicts continuous affective states from music, while a CLIP based visual VA head estimates emotions from images. These components jointly enforce semantic and emotional alignment between music and synthesized images. Experiments on curated music image pairs show that MESA MIG outperforms caption only and single agent baselines in aesthetic quality, semantic consistency, and VA alignment, and achieves competitive emotion regression performance compared with state of the art music and image emotion models.
format Preprint
id arxiv_https___arxiv_org_abs_2512_23320
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi Agents Semantic Emotion Aligned Music to Image Generation with Music Derived Captions
Shi, Junchang
Li, Gang
Multimedia
When people listen to music, they often experience rich visual imagery. We aim to externalize this inner imagery by generating images conditioned on music. We propose MESA MIG, a multi agent semantic and emotion aligned framework that first produces structured music captions and then refines them with cooperating agents specializing in scene, motion, style, color, and composition. In parallel, a Valence Arousal regression head predicts continuous affective states from music, while a CLIP based visual VA head estimates emotions from images. These components jointly enforce semantic and emotional alignment between music and synthesized images. Experiments on curated music image pairs show that MESA MIG outperforms caption only and single agent baselines in aesthetic quality, semantic consistency, and VA alignment, and achieves competitive emotion regression performance compared with state of the art music and image emotion models.
title Multi Agents Semantic Emotion Aligned Music to Image Generation with Music Derived Captions
topic Multimedia
url https://arxiv.org/abs/2512.23320