Do Joint Language-Audio Embeddings Encode Perceptual Timbre Semantics?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Deng, Qixin, Pardo, Bryan, Pappas, Thrasyvoulos N
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909849744310272
author Deng, Qixin
Pardo, Bryan
Pappas, Thrasyvoulos N
author_facet Deng, Qixin
Pardo, Bryan
Pappas, Thrasyvoulos N
contents Understanding and modeling the relationship between language and sound is critical for applications such as music information retrieval,text-guided music generation, and audio captioning. Central to these tasks is the use of joint language-audio embedding spaces, which map textual descriptions and auditory content into a shared embedding space. While multimodal embedding models such as MS-CLAP, LAION-CLAP, and MuQ-MuLan have shown strong performance in aligning language and audio, their correspondence to human perception of timbre, a multifaceted attribute encompassing qualities such as brightness, roughness, and warmth, remains underexplored. In this paper, we evaluate the above three joint language-audio embedding models on their ability to capture perceptual dimensions of timbre. Our findings show that LAION-CLAP consistently provides the most reliable alignment with human-perceived timbre semantics across both instrumental sounds and audio effects.
format Preprint
id arxiv_https___arxiv_org_abs_2510_14249
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Do Joint Language-Audio Embeddings Encode Perceptual Timbre Semantics?
Deng, Qixin
Pardo, Bryan
Pappas, Thrasyvoulos N
Sound
Artificial Intelligence
Audio and Speech Processing
Understanding and modeling the relationship between language and sound is critical for applications such as music information retrieval,text-guided music generation, and audio captioning. Central to these tasks is the use of joint language-audio embedding spaces, which map textual descriptions and auditory content into a shared embedding space. While multimodal embedding models such as MS-CLAP, LAION-CLAP, and MuQ-MuLan have shown strong performance in aligning language and audio, their correspondence to human perception of timbre, a multifaceted attribute encompassing qualities such as brightness, roughness, and warmth, remains underexplored. In this paper, we evaluate the above three joint language-audio embedding models on their ability to capture perceptual dimensions of timbre. Our findings show that LAION-CLAP consistently provides the most reliable alignment with human-perceived timbre semantics across both instrumental sounds and audio effects.
title Do Joint Language-Audio Embeddings Encode Perceptual Timbre Semantics?
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2510.14249