Revisiting MFCCs: Evidence for Spectral-Prosodic Coupling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bezerra, Vitor Magno de O. S., Bastos, Gabriel F. A., Montalvão, Jugurta
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909829579145216
author Bezerra, Vitor Magno de O. S.
Bastos, Gabriel F. A.
Montalvão, Jugurta
author_facet Bezerra, Vitor Magno de O. S.
Bastos, Gabriel F. A.
Montalvão, Jugurta
contents Mel-frequency cepstral coefficients (MFCCs) are an important feature in speech processing. A deeper understanding of their properties can contribute to the work that is being done with both classical and deep learning models. This study challenges the long-held assumption that MFCCs lack relevant temporal information by investigating their relationship with speech prosody. Using a null hypothesis significance testing framework, a systematic assessment is made about the statistical independence between MFCCs and the three prosodic features: energy, fundamental frequency (F0), and voicing. The results demonstrate that it is statistically implausible that the MFCCs are independent of any of these three prosodic features. This finding suggests that MFCCs inherently carry valuable prosodic information, which can inform the design of future models in speech analysis and recognition.
format Preprint
id arxiv_https___arxiv_org_abs_2510_05922
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Revisiting MFCCs: Evidence for Spectral-Prosodic Coupling
Bezerra, Vitor Magno de O. S.
Bastos, Gabriel F. A.
Montalvão, Jugurta
Audio and Speech Processing
Mel-frequency cepstral coefficients (MFCCs) are an important feature in speech processing. A deeper understanding of their properties can contribute to the work that is being done with both classical and deep learning models. This study challenges the long-held assumption that MFCCs lack relevant temporal information by investigating their relationship with speech prosody. Using a null hypothesis significance testing framework, a systematic assessment is made about the statistical independence between MFCCs and the three prosodic features: energy, fundamental frequency (F0), and voicing. The results demonstrate that it is statistically implausible that the MFCCs are independent of any of these three prosodic features. This finding suggests that MFCCs inherently carry valuable prosodic information, which can inform the design of future models in speech analysis and recognition.
title Revisiting MFCCs: Evidence for Spectral-Prosodic Coupling
topic Audio and Speech Processing
url https://arxiv.org/abs/2510.05922