Emotion-Aligned Contrastive Learning Between Images and Music

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Stewart, Shanti, Avramidis, Kleanthis, Feng, Tiantian, Narayanan, Shrikanth
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913600275218432
author Stewart, Shanti
Avramidis, Kleanthis
Feng, Tiantian
Narayanan, Shrikanth
author_facet Stewart, Shanti
Avramidis, Kleanthis
Feng, Tiantian
Narayanan, Shrikanth
contents Traditional music search engines rely on retrieval methods that match natural language queries with music metadata. There have been increasing efforts to expand retrieval methods to consider the audio characteristics of music itself, using queries of various modalities including text, video, and speech. While most approaches aim to match general music semantics to the input queries, only a few focus on affective qualities. In this work, we address the task of retrieving emotionally-relevant music from image queries by learning an affective alignment between images and music audio. Our approach focuses on learning an emotion-aligned joint embedding space between images and music. This embedding space is learned via emotion-supervised contrastive learning, using an adapted cross-modal version of the SupCon loss. We evaluate the joint embeddings through cross-modal retrieval tasks (image-to-music and music-to-image) based on emotion labels. Furthermore, we investigate the generalizability of the learned music embeddings via automatic music tagging. Our experiments show that the proposed approach successfully aligns images and music, and that the learned embedding space is effective for cross-modal retrieval applications.
format Preprint
id arxiv_https___arxiv_org_abs_2308_12610
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Emotion-Aligned Contrastive Learning Between Images and Music
Stewart, Shanti
Avramidis, Kleanthis
Feng, Tiantian
Narayanan, Shrikanth
Multimedia
Sound
Audio and Speech Processing
Traditional music search engines rely on retrieval methods that match natural language queries with music metadata. There have been increasing efforts to expand retrieval methods to consider the audio characteristics of music itself, using queries of various modalities including text, video, and speech. While most approaches aim to match general music semantics to the input queries, only a few focus on affective qualities. In this work, we address the task of retrieving emotionally-relevant music from image queries by learning an affective alignment between images and music audio. Our approach focuses on learning an emotion-aligned joint embedding space between images and music. This embedding space is learned via emotion-supervised contrastive learning, using an adapted cross-modal version of the SupCon loss. We evaluate the joint embeddings through cross-modal retrieval tasks (image-to-music and music-to-image) based on emotion labels. Furthermore, we investigate the generalizability of the learned music embeddings via automatic music tagging. Our experiments show that the proposed approach successfully aligns images and music, and that the learned embedding space is effective for cross-modal retrieval applications.
title Emotion-Aligned Contrastive Learning Between Images and Music
topic Multimedia
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2308.12610