Exploring System Adaptations For Minimum Latency Real-Time Piano Transcription

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Patricia, Peter, Silvan David, Schlüter, Jan, Widmer, Gerhard
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916942552498176
author Hu, Patricia
Peter, Silvan David
Schlüter, Jan
Widmer, Gerhard
author_facet Hu, Patricia
Peter, Silvan David
Schlüter, Jan
Widmer, Gerhard
contents Advances in neural network design and the availability of large-scale labeled datasets have driven major improvements in piano transcription. Existing approaches target either offline applications, with no restrictions on computational demands, or online transcription, with delays of 128-320 ms. However, most real-time musical applications require latencies below 30 ms. In this work, we investigate whether and how the current state-of-the-art online transcription model can be adapted for real-time piano transcription. Specifically, we eliminate all non-causal processing, and reduce computational load through shared computations across core model components and variations in model size. Additionally, we explore different pre- and postprocessing strategies, and related label encoding schemes, and discuss their suitability for real-time transcription. Evaluating the adaptions on the MAESTRO dataset, we find a drop in transcription accuracy due to strictly causal processing as well as a tradeoff between the preprocessing latency and prediction accuracy. We release our system as a baseline to support researchers in designing models towards minimum latency real-time transcription.
format Preprint
id arxiv_https___arxiv_org_abs_2509_07586
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring System Adaptations For Minimum Latency Real-Time Piano Transcription
Hu, Patricia
Peter, Silvan David
Schlüter, Jan
Widmer, Gerhard
Audio and Speech Processing
Machine Learning
Sound
Advances in neural network design and the availability of large-scale labeled datasets have driven major improvements in piano transcription. Existing approaches target either offline applications, with no restrictions on computational demands, or online transcription, with delays of 128-320 ms. However, most real-time musical applications require latencies below 30 ms. In this work, we investigate whether and how the current state-of-the-art online transcription model can be adapted for real-time piano transcription. Specifically, we eliminate all non-causal processing, and reduce computational load through shared computations across core model components and variations in model size. Additionally, we explore different pre- and postprocessing strategies, and related label encoding schemes, and discuss their suitability for real-time transcription. Evaluating the adaptions on the MAESTRO dataset, we find a drop in transcription accuracy due to strictly causal processing as well as a tradeoff between the preprocessing latency and prediction accuracy. We release our system as a baseline to support researchers in designing models towards minimum latency real-time transcription.
title Exploring System Adaptations For Minimum Latency Real-Time Piano Transcription
topic Audio and Speech Processing
Machine Learning
Sound
url https://arxiv.org/abs/2509.07586