Harmonics to the Rescue: Why Voiced Speech is Not a Wss Process

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Bologni, Giovanni, Heusdens, Richard, Hendriks, Richard C.
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908448875085824
author Bologni, Giovanni
Heusdens, Richard
Hendriks, Richard C.
author_facet Bologni, Giovanni
Heusdens, Richard
Hendriks, Richard C.
contents Speech processing algorithms often rely on statistical knowledge of the underlying process. Despite many years of research, however, the debate on the most appropriate statistical model for speech still continues. Speech is commonly modeled as a wide-sense stationary (WSS) process. However, the use of the WSS model for spectrally correlated processes is fundamentally wrong, as WSS implies spectral uncorrelation. In this paper, we demonstrate that voiced speech can be more accurately represented as a cyclostationary (CS) process. By employing the CS rather than the WSS model for processes that are inherently correlated across frequency, it is possible to improve the estimation of cross-power spectral densities (PSDs), source separation, and beamforming. We illustrate how the correlation between harmonic frequencies of CS processes can enhance system identification, and validate our findings using both simulated and real speech data.
format Preprint
id arxiv_https___arxiv_org_abs_2507_10176
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Harmonics to the Rescue: Why Voiced Speech is Not a Wss Process
Bologni, Giovanni
Heusdens, Richard
Hendriks, Richard C.
Audio and Speech Processing
Signal Processing
Speech processing algorithms often rely on statistical knowledge of the underlying process. Despite many years of research, however, the debate on the most appropriate statistical model for speech still continues. Speech is commonly modeled as a wide-sense stationary (WSS) process. However, the use of the WSS model for spectrally correlated processes is fundamentally wrong, as WSS implies spectral uncorrelation. In this paper, we demonstrate that voiced speech can be more accurately represented as a cyclostationary (CS) process. By employing the CS rather than the WSS model for processes that are inherently correlated across frequency, it is possible to improve the estimation of cross-power spectral densities (PSDs), source separation, and beamforming. We illustrate how the correlation between harmonic frequencies of CS processes can enhance system identification, and validate our findings using both simulated and real speech data.
title Harmonics to the Rescue: Why Voiced Speech is Not a Wss Process
topic Audio and Speech Processing
Signal Processing
url https://arxiv.org/abs/2507.10176