Are we describing the same sound? An analysis of word embedding spaces of expressive piano performance

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Peter, Silvan David, Chowdhury, Shreyan, Cancino-Chacón, Carlos Eduardo, Widmer, Gerhard
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909063317553152
author Peter, Silvan David
Chowdhury, Shreyan
Cancino-Chacón, Carlos Eduardo
Widmer, Gerhard
author_facet Peter, Silvan David
Chowdhury, Shreyan
Cancino-Chacón, Carlos Eduardo
Widmer, Gerhard
contents Semantic embeddings play a crucial role in natural language-based information retrieval. Embedding models represent words and contexts as vectors whose spatial configuration is derived from the distribution of words in large text corpora. While such representations are generally very powerful, they might fail to account for fine-grained domain-specific nuances. In this article, we investigate this uncertainty for the domain of characterizations of expressive piano performance. Using a music research dataset of free text performance characterizations and a follow-up study sorting the annotations into clusters, we derive a ground truth for a domain-specific semantic similarity structure. We test five embedding models and their similarity structure for correspondence with the ground truth. We further assess the effects of contextualizing prompts, hubness reduction, cross-modal similarity, and k-means clustering. The quality of embedding models shows great variability with respect to this task; more general models perform better than domain-adapted ones and the best model configurations reach human-level agreement.
format Preprint
id arxiv_https___arxiv_org_abs_2401_02979
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Are we describing the same sound? An analysis of word embedding spaces of expressive piano performance
Peter, Silvan David
Chowdhury, Shreyan
Cancino-Chacón, Carlos Eduardo
Widmer, Gerhard
Computation and Language
Artificial Intelligence
Information Retrieval
Semantic embeddings play a crucial role in natural language-based information retrieval. Embedding models represent words and contexts as vectors whose spatial configuration is derived from the distribution of words in large text corpora. While such representations are generally very powerful, they might fail to account for fine-grained domain-specific nuances. In this article, we investigate this uncertainty for the domain of characterizations of expressive piano performance. Using a music research dataset of free text performance characterizations and a follow-up study sorting the annotations into clusters, we derive a ground truth for a domain-specific semantic similarity structure. We test five embedding models and their similarity structure for correspondence with the ground truth. We further assess the effects of contextualizing prompts, hubness reduction, cross-modal similarity, and k-means clustering. The quality of embedding models shows great variability with respect to this task; more general models perform better than domain-adapted ones and the best model configurations reach human-level agreement.
title Are we describing the same sound? An analysis of word embedding spaces of expressive piano performance
topic Computation and Language
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2401.02979