Real-Time Streaming Mel Vocoding with Generative Flow Matching

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Welker, Simon, Peer, Tal, Gerkmann, Timo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911161908199424
author Welker, Simon
Peer, Tal
Gerkmann, Timo
author_facet Welker, Simon
Peer, Tal
Gerkmann, Timo
contents The task of Mel vocoding, i.e., the inversion of a Mel magnitude spectrogram to an audio waveform, is still a key component in many text-to-speech (TTS) systems today. Based on generative flow matching, our prior work on generative STFT phase retrieval (DiffPhase), and the pseudoinverse operator of the Mel filterbank, we develop MelFlow, a streaming-capable generative Mel vocoder for speech sampled at 16 kHz with an algorithmic latency of only 32 ms and a total latency of 48 ms. We show real-time streaming capability at this latency not only in theory, but in practice on a consumer laptop GPU. Furthermore, we show that our model achieves substantially better PESQ and SI-SDR values compared to well-established not streaming-capable baselines for Mel vocoding including HiFi-GAN.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15085
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Real-Time Streaming Mel Vocoding with Generative Flow Matching
Welker, Simon
Peer, Tal
Gerkmann, Timo
Audio and Speech Processing
Machine Learning
Sound
Signal Processing
The task of Mel vocoding, i.e., the inversion of a Mel magnitude spectrogram to an audio waveform, is still a key component in many text-to-speech (TTS) systems today. Based on generative flow matching, our prior work on generative STFT phase retrieval (DiffPhase), and the pseudoinverse operator of the Mel filterbank, we develop MelFlow, a streaming-capable generative Mel vocoder for speech sampled at 16 kHz with an algorithmic latency of only 32 ms and a total latency of 48 ms. We show real-time streaming capability at this latency not only in theory, but in practice on a consumer laptop GPU. Furthermore, we show that our model achieves substantially better PESQ and SI-SDR values compared to well-established not streaming-capable baselines for Mel vocoding including HiFi-GAN.
title Real-Time Streaming Mel Vocoding with Generative Flow Matching
topic Audio and Speech Processing
Machine Learning
Sound
Signal Processing
url https://arxiv.org/abs/2509.15085