Saved in:
Bibliographic Details
Main Authors: De Cristofaro, Domenico, Vitale, Vincenzo Norman, Vietti, Alessandro
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2508.17914
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912553027764224
author De Cristofaro, Domenico
Vitale, Vincenzo Norman
Vietti, Alessandro
author_facet De Cristofaro, Domenico
Vitale, Vincenzo Norman
Vietti, Alessandro
contents Automatic Speech Recognition has advanced with self-supervised learning, enabling feature extraction directly from raw audio. In Wav2Vec, a CNN first transforms audio into feature vectors before the transformer processes them. This study examines CNN-extracted information for monophthong vowels using the TIMIT corpus. We compare MFCCs, MFCCs with formants, and CNN activations by training SVM classifiers for front-back vowel identification, assessing their classification accuracy to evaluate phonetic representation.
format Preprint
id arxiv_https___arxiv_org_abs_2508_17914
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating the Representation of Vowels in Wav2Vec Feature Extractor: A Layer-Wise Analysis Using MFCCs
De Cristofaro, Domenico
Vitale, Vincenzo Norman
Vietti, Alessandro
Computation and Language
Automatic Speech Recognition has advanced with self-supervised learning, enabling feature extraction directly from raw audio. In Wav2Vec, a CNN first transforms audio into feature vectors before the transformer processes them. This study examines CNN-extracted information for monophthong vowels using the TIMIT corpus. We compare MFCCs, MFCCs with formants, and CNN activations by training SVM classifiers for front-back vowel identification, assessing their classification accuracy to evaluate phonetic representation.
title Evaluating the Representation of Vowels in Wav2Vec Feature Extractor: A Layer-Wise Analysis Using MFCCs
topic Computation and Language
url https://arxiv.org/abs/2508.17914