Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency Streaming ASR

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Longhao, Li, Yangze, Xue, Hongfei, Liu, Jie, Fang, Shuai, Wang, Kai, Xie, Lei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912399999631360
author Li, Longhao
Li, Yangze
Xue, Hongfei
Liu, Jie
Fang, Shuai
Wang, Kai
Xie, Lei
author_facet Li, Longhao
Li, Yangze
Xue, Hongfei
Liu, Jie
Fang, Shuai
Wang, Kai
Xie, Lei
contents CTC-based streaming ASR has gained significant attention in real-world applications but faces two main challenges: accuracy degradation in small chunks and token emission latency. To mitigate these challenges, we propose Delayed-KD, which applies delayed knowledge distillation on CTC posterior probabilities from a non-streaming to a streaming model. Specifically, with a tiny chunk size, we introduce a Temporal Alignment Buffer (TAB) that defines a relative delay range compared to the non-streaming teacher model to align CTC outputs and mitigate non-blank token mismatches. Additionally, TAB enables fine-grained control over token emission delay. Experiments on 178-hour AISHELL-1 and 10,000-hour WenetSpeech Mandarin datasets show consistent superiority of Delayed-KD. Impressively, Delayed-KD at 40 ms latency achieves a lower character error rate (CER) of 5.42% on AISHELL-1, comparable to the competitive U2++ model running at 320 ms latency.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22069
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency Streaming ASR
Li, Longhao
Li, Yangze
Xue, Hongfei
Liu, Jie
Fang, Shuai
Wang, Kai
Xie, Lei
Sound
Audio and Speech Processing
CTC-based streaming ASR has gained significant attention in real-world applications but faces two main challenges: accuracy degradation in small chunks and token emission latency. To mitigate these challenges, we propose Delayed-KD, which applies delayed knowledge distillation on CTC posterior probabilities from a non-streaming to a streaming model. Specifically, with a tiny chunk size, we introduce a Temporal Alignment Buffer (TAB) that defines a relative delay range compared to the non-streaming teacher model to align CTC outputs and mitigate non-blank token mismatches. Additionally, TAB enables fine-grained control over token emission delay. Experiments on 178-hour AISHELL-1 and 10,000-hour WenetSpeech Mandarin datasets show consistent superiority of Delayed-KD. Impressively, Delayed-KD at 40 ms latency achieves a lower character error rate (CER) of 5.42% on AISHELL-1, comparable to the competitive U2++ model running at 320 ms latency.
title Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency Streaming ASR
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.22069