LCS-CTC: Leveraging Soft Alignments to Enhance Phonetic Transcription Robustness

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ye, Zongli, Lian, Jiachen, Gupta, Akshaj, Zhou, Xuanru, Li, Haodong, Patel, Krish, Park, Hwi Joo, Zhou, Dingkun, Guo, Chenxu, Li, Shuhe, Wang, Sam, Zhou, Iris, Cho, Cheol Jun, Ezzes, Zoe, Vonk, Jet M. J., Morin, Brittany T., Bogley, Rian, Wauters, Lisa, Miller, Zachary A., Gorno-Tempini, Maria Luisa, Anumanchipalli, Gopala
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912536896471040
author Ye, Zongli
Lian, Jiachen
Gupta, Akshaj
Zhou, Xuanru
Li, Haodong
Patel, Krish
Park, Hwi Joo
Zhou, Dingkun
Guo, Chenxu
Li, Shuhe
Wang, Sam
Zhou, Iris
Cho, Cheol Jun
Ezzes, Zoe
Vonk, Jet M. J.
Morin, Brittany T.
Bogley, Rian
Wauters, Lisa
Miller, Zachary A.
Gorno-Tempini, Maria Luisa
Anumanchipalli, Gopala
author_facet Ye, Zongli
Lian, Jiachen
Gupta, Akshaj
Zhou, Xuanru
Li, Haodong
Patel, Krish
Park, Hwi Joo
Zhou, Dingkun
Guo, Chenxu
Li, Shuhe
Wang, Sam
Zhou, Iris
Cho, Cheol Jun
Ezzes, Zoe
Vonk, Jet M. J.
Morin, Brittany T.
Bogley, Rian
Wauters, Lisa
Miller, Zachary A.
Gorno-Tempini, Maria Luisa
Anumanchipalli, Gopala
contents Phonetic speech transcription is crucial for fine-grained linguistic analysis and downstream speech applications. While Connectionist Temporal Classification (CTC) is a widely used approach for such tasks due to its efficiency, it often falls short in recognition performance, especially under unclear and nonfluent speech. In this work, we propose LCS-CTC, a two-stage framework for phoneme-level speech recognition that combines a similarity-aware local alignment algorithm with a constrained CTC training objective. By predicting fine-grained frame-phoneme cost matrices and applying a modified Longest Common Subsequence (LCS) algorithm, our method identifies high-confidence alignment zones which are used to constrain the CTC decoding path space, thereby reducing overfitting and improving generalization ability, which enables both robust recognition and text-free forced alignment. Experiments on both LibriSpeech and PPA demonstrate that LCS-CTC consistently outperforms vanilla CTC baselines, suggesting its potential to unify phoneme modeling across fluent and non-fluent speech.
format Preprint
id arxiv_https___arxiv_org_abs_2508_03937
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LCS-CTC: Leveraging Soft Alignments to Enhance Phonetic Transcription Robustness
Ye, Zongli
Lian, Jiachen
Gupta, Akshaj
Zhou, Xuanru
Li, Haodong
Patel, Krish
Park, Hwi Joo
Zhou, Dingkun
Guo, Chenxu
Li, Shuhe
Wang, Sam
Zhou, Iris
Cho, Cheol Jun
Ezzes, Zoe
Vonk, Jet M. J.
Morin, Brittany T.
Bogley, Rian
Wauters, Lisa
Miller, Zachary A.
Gorno-Tempini, Maria Luisa
Anumanchipalli, Gopala
Audio and Speech Processing
Phonetic speech transcription is crucial for fine-grained linguistic analysis and downstream speech applications. While Connectionist Temporal Classification (CTC) is a widely used approach for such tasks due to its efficiency, it often falls short in recognition performance, especially under unclear and nonfluent speech. In this work, we propose LCS-CTC, a two-stage framework for phoneme-level speech recognition that combines a similarity-aware local alignment algorithm with a constrained CTC training objective. By predicting fine-grained frame-phoneme cost matrices and applying a modified Longest Common Subsequence (LCS) algorithm, our method identifies high-confidence alignment zones which are used to constrain the CTC decoding path space, thereby reducing overfitting and improving generalization ability, which enables both robust recognition and text-free forced alignment. Experiments on both LibriSpeech and PPA demonstrate that LCS-CTC consistently outperforms vanilla CTC baselines, suggesting its potential to unify phoneme modeling across fluent and non-fluent speech.
title LCS-CTC: Leveraging Soft Alignments to Enhance Phonetic Transcription Robustness
topic Audio and Speech Processing
url https://arxiv.org/abs/2508.03937