Sounding Out Reconstruction Error-Based Evaluation of Generative Models of Expressive Performance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Peter, Silvan David, Cancino-Chacón, Carlos Eduardo, Karystinaios, Emmanouil, Widmer, Gerhard
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929195774377984
author Peter, Silvan David
Cancino-Chacón, Carlos Eduardo
Karystinaios, Emmanouil
Widmer, Gerhard
author_facet Peter, Silvan David
Cancino-Chacón, Carlos Eduardo
Karystinaios, Emmanouil
Widmer, Gerhard
contents Generative models of expressive piano performance are usually assessed by comparing their predictions to a reference human performance. A generative algorithm is taken to be better than competing ones if it produces performances that are closer to a human reference performance. However, expert human performers can (and do) interpret music in different ways, making for different possible references, and quantitative closeness is not necessarily aligned with perceptual similarity, raising concerns about the validity of this evaluation approach. In this work, we present a number of experiments that shed light on this problem. Using precisely measured high-quality performances of classical piano music, we carry out a listening test indicating that listeners can sometimes perceive subtle performance difference that go unnoticed under quantitative evaluation. We further present tests that indicate that such evaluation frameworks show a lot of variability in reliability and validity across different reference performances and pieces. We discuss these results and their implications for quantitative evaluation, and hope to foster a critical appreciation of the uncertainties involved in quantitative assessments of such performances within the wider music information retrieval (MIR) community.
format Preprint
id arxiv_https___arxiv_org_abs_2401_00471
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Sounding Out Reconstruction Error-Based Evaluation of Generative Models of Expressive Performance
Peter, Silvan David
Cancino-Chacón, Carlos Eduardo
Karystinaios, Emmanouil
Widmer, Gerhard
Sound
Audio and Speech Processing
Generative models of expressive piano performance are usually assessed by comparing their predictions to a reference human performance. A generative algorithm is taken to be better than competing ones if it produces performances that are closer to a human reference performance. However, expert human performers can (and do) interpret music in different ways, making for different possible references, and quantitative closeness is not necessarily aligned with perceptual similarity, raising concerns about the validity of this evaluation approach. In this work, we present a number of experiments that shed light on this problem. Using precisely measured high-quality performances of classical piano music, we carry out a listening test indicating that listeners can sometimes perceive subtle performance difference that go unnoticed under quantitative evaluation. We further present tests that indicate that such evaluation frameworks show a lot of variability in reliability and validity across different reference performances and pieces. We discuss these results and their implications for quantitative evaluation, and hope to foster a critical appreciation of the uncertainties involved in quantitative assessments of such performances within the wider music information retrieval (MIR) community.
title Sounding Out Reconstruction Error-Based Evaluation of Generative Models of Expressive Performance
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2401.00471