LCS-CTC: Leveraging Soft Alignments to Enhance Phonetic Transcription Robustness
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866912536896471040 |
|---|---|
| author | Ye, Zongli Lian, Jiachen Gupta, Akshaj Zhou, Xuanru Li, Haodong Patel, Krish Park, Hwi Joo Zhou, Dingkun Guo, Chenxu Li, Shuhe Wang, Sam Zhou, Iris Cho, Cheol Jun Ezzes, Zoe Vonk, Jet M. J. Morin, Brittany T. Bogley, Rian Wauters, Lisa Miller, Zachary A. Gorno-Tempini, Maria Luisa Anumanchipalli, Gopala |
| author_facet | Ye, Zongli Lian, Jiachen Gupta, Akshaj Zhou, Xuanru Li, Haodong Patel, Krish Park, Hwi Joo Zhou, Dingkun Guo, Chenxu Li, Shuhe Wang, Sam Zhou, Iris Cho, Cheol Jun Ezzes, Zoe Vonk, Jet M. J. Morin, Brittany T. Bogley, Rian Wauters, Lisa Miller, Zachary A. Gorno-Tempini, Maria Luisa Anumanchipalli, Gopala |
| contents | Phonetic speech transcription is crucial for fine-grained linguistic analysis and downstream speech applications. While Connectionist Temporal Classification (CTC) is a widely used approach for such tasks due to its efficiency, it often falls short in recognition performance, especially under unclear and nonfluent speech. In this work, we propose LCS-CTC, a two-stage framework for phoneme-level speech recognition that combines a similarity-aware local alignment algorithm with a constrained CTC training objective. By predicting fine-grained frame-phoneme cost matrices and applying a modified Longest Common Subsequence (LCS) algorithm, our method identifies high-confidence alignment zones which are used to constrain the CTC decoding path space, thereby reducing overfitting and improving generalization ability, which enables both robust recognition and text-free forced alignment. Experiments on both LibriSpeech and PPA demonstrate that LCS-CTC consistently outperforms vanilla CTC baselines, suggesting its potential to unify phoneme modeling across fluent and non-fluent speech. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_03937 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | LCS-CTC: Leveraging Soft Alignments to Enhance Phonetic Transcription Robustness Ye, Zongli Lian, Jiachen Gupta, Akshaj Zhou, Xuanru Li, Haodong Patel, Krish Park, Hwi Joo Zhou, Dingkun Guo, Chenxu Li, Shuhe Wang, Sam Zhou, Iris Cho, Cheol Jun Ezzes, Zoe Vonk, Jet M. J. Morin, Brittany T. Bogley, Rian Wauters, Lisa Miller, Zachary A. Gorno-Tempini, Maria Luisa Anumanchipalli, Gopala Audio and Speech Processing Phonetic speech transcription is crucial for fine-grained linguistic analysis and downstream speech applications. While Connectionist Temporal Classification (CTC) is a widely used approach for such tasks due to its efficiency, it often falls short in recognition performance, especially under unclear and nonfluent speech. In this work, we propose LCS-CTC, a two-stage framework for phoneme-level speech recognition that combines a similarity-aware local alignment algorithm with a constrained CTC training objective. By predicting fine-grained frame-phoneme cost matrices and applying a modified Longest Common Subsequence (LCS) algorithm, our method identifies high-confidence alignment zones which are used to constrain the CTC decoding path space, thereby reducing overfitting and improving generalization ability, which enables both robust recognition and text-free forced alignment. Experiments on both LibriSpeech and PPA demonstrate that LCS-CTC consistently outperforms vanilla CTC baselines, suggesting its potential to unify phoneme modeling across fluent and non-fluent speech. |
| title | LCS-CTC: Leveraging Soft Alignments to Enhance Phonetic Transcription Robustness |
| topic | Audio and Speech Processing |
| url | https://arxiv.org/abs/2508.03937 |