Automatic classification of stop realisation with wav2vec2.0

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tanner, James, Sonderegger, Morgan, Stuart-Smith, Jane, Mielke, Jeff, Kendall, Tyler
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912403370803200
author Tanner, James
Sonderegger, Morgan
Stuart-Smith, Jane
Mielke, Jeff
Kendall, Tyler
author_facet Tanner, James
Sonderegger, Morgan
Stuart-Smith, Jane
Mielke, Jeff
Kendall, Tyler
contents Modern phonetic research regularly makes use of automatic tools for the annotation of speech data, however few tools exist for the annotation of many variable phonetic phenomena. At the same time, pre-trained self-supervised models, such as wav2vec2.0, have been shown to perform well at speech classification tasks and latently encode fine-grained phonetic information. We demonstrate that wav2vec2.0 models can be trained to automatically classify stop burst presence with high accuracy in both English and Japanese, robust across both finely-curated and unprepared speech corpora. Patterns of variability in stop realisation are replicated with the automatic annotations, and closely follow those of manual annotations. These results demonstrate the potential of pre-trained speech models as tools for the automatic annotation and processing of speech corpus data, enabling researchers to 'scale-up' the scope of phonetic research with relative ease.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23688
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Automatic classification of stop realisation with wav2vec2.0
Tanner, James
Sonderegger, Morgan
Stuart-Smith, Jane
Mielke, Jeff
Kendall, Tyler
Computation and Language
Sound
Audio and Speech Processing
Modern phonetic research regularly makes use of automatic tools for the annotation of speech data, however few tools exist for the annotation of many variable phonetic phenomena. At the same time, pre-trained self-supervised models, such as wav2vec2.0, have been shown to perform well at speech classification tasks and latently encode fine-grained phonetic information. We demonstrate that wav2vec2.0 models can be trained to automatically classify stop burst presence with high accuracy in both English and Japanese, robust across both finely-curated and unprepared speech corpora. Patterns of variability in stop realisation are replicated with the automatic annotations, and closely follow those of manual annotations. These results demonstrate the potential of pre-trained speech models as tools for the automatic annotation and processing of speech corpus data, enabling researchers to 'scale-up' the scope of phonetic research with relative ease.
title Automatic classification of stop realisation with wav2vec2.0
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.23688