Graph Connectionist Temporal Classification for Phoneme Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Grafé, Henry, Van hamme, Hugo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914025595469824
author Grafé, Henry
Van hamme, Hugo
author_facet Grafé, Henry
Van hamme, Hugo
contents Automatic Phoneme Recognition (APR) systems are often trained using pseudo phoneme-level annotations generated from text through Grapheme-to-Phoneme (G2P) systems. These G2P systems frequently output multiple possible pronunciations per word, but the standard Connectionist Temporal Classification (CTC) loss cannot account for such ambiguity during training. In this work, we adapt Graph Temporal Classification (GTC) to the APR setting. GTC enables training from a graph of alternative phoneme sequences, allowing the model to consider multiple pronunciations per word as valid supervision. Our experiments on English and Dutch data sets show that incorporating multiple pronunciations per word into the training loss consistently improves phoneme error rates compared to a baseline trained with CTC. These results suggest that integrating pronunciation variation into the loss function is a promising strategy for training APR systems from noisy G2P-based supervision.
format Preprint
id arxiv_https___arxiv_org_abs_2509_05399
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Graph Connectionist Temporal Classification for Phoneme Recognition
Grafé, Henry
Van hamme, Hugo
Audio and Speech Processing
Artificial Intelligence
68T10, 68T07
I.2.7; I.5.4; H.5.1
Automatic Phoneme Recognition (APR) systems are often trained using pseudo phoneme-level annotations generated from text through Grapheme-to-Phoneme (G2P) systems. These G2P systems frequently output multiple possible pronunciations per word, but the standard Connectionist Temporal Classification (CTC) loss cannot account for such ambiguity during training. In this work, we adapt Graph Temporal Classification (GTC) to the APR setting. GTC enables training from a graph of alternative phoneme sequences, allowing the model to consider multiple pronunciations per word as valid supervision. Our experiments on English and Dutch data sets show that incorporating multiple pronunciations per word into the training loss consistently improves phoneme error rates compared to a baseline trained with CTC. These results suggest that integrating pronunciation variation into the loss function is a promising strategy for training APR systems from noisy G2P-based supervision.
title Graph Connectionist Temporal Classification for Phoneme Recognition
topic Audio and Speech Processing
Artificial Intelligence
68T10, 68T07
I.2.7; I.5.4; H.5.1
url https://arxiv.org/abs/2509.05399