Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency Streaming ASR
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912399999631360 |
|---|---|
| author | Li, Longhao Li, Yangze Xue, Hongfei Liu, Jie Fang, Shuai Wang, Kai Xie, Lei |
| author_facet | Li, Longhao Li, Yangze Xue, Hongfei Liu, Jie Fang, Shuai Wang, Kai Xie, Lei |
| contents | CTC-based streaming ASR has gained significant attention in real-world applications but faces two main challenges: accuracy degradation in small chunks and token emission latency. To mitigate these challenges, we propose Delayed-KD, which applies delayed knowledge distillation on CTC posterior probabilities from a non-streaming to a streaming model. Specifically, with a tiny chunk size, we introduce a Temporal Alignment Buffer (TAB) that defines a relative delay range compared to the non-streaming teacher model to align CTC outputs and mitigate non-blank token mismatches. Additionally, TAB enables fine-grained control over token emission delay. Experiments on 178-hour AISHELL-1 and 10,000-hour WenetSpeech Mandarin datasets show consistent superiority of Delayed-KD. Impressively, Delayed-KD at 40 ms latency achieves a lower character error rate (CER) of 5.42% on AISHELL-1, comparable to the competitive U2++ model running at 320 ms latency. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_22069 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency Streaming ASR Li, Longhao Li, Yangze Xue, Hongfei Liu, Jie Fang, Shuai Wang, Kai Xie, Lei Sound Audio and Speech Processing CTC-based streaming ASR has gained significant attention in real-world applications but faces two main challenges: accuracy degradation in small chunks and token emission latency. To mitigate these challenges, we propose Delayed-KD, which applies delayed knowledge distillation on CTC posterior probabilities from a non-streaming to a streaming model. Specifically, with a tiny chunk size, we introduce a Temporal Alignment Buffer (TAB) that defines a relative delay range compared to the non-streaming teacher model to align CTC outputs and mitigate non-blank token mismatches. Additionally, TAB enables fine-grained control over token emission delay. Experiments on 178-hour AISHELL-1 and 10,000-hour WenetSpeech Mandarin datasets show consistent superiority of Delayed-KD. Impressively, Delayed-KD at 40 ms latency achieves a lower character error rate (CER) of 5.42% on AISHELL-1, comparable to the competitive U2++ model running at 320 ms latency. |
| title | Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency Streaming ASR |
| topic | Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2505.22069 |