M-CIF: Multi-Scale Alignment For CIF-Based Non-Autoregressive ASR
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918171220377600 |
|---|---|
| author | Mao, Ruixiang Ma, Xiangnan Yang, Qing Zhu, Ziming Qiao, Yucheng Ge, Yuan Xiao, Tong Gao, Shengxiang Yu, Zhengtao Zhu, Jingbo |
| author_facet | Mao, Ruixiang Ma, Xiangnan Yang, Qing Zhu, Ziming Qiao, Yucheng Ge, Yuan Xiao, Tong Gao, Shengxiang Yu, Zhengtao Zhu, Jingbo |
| contents | The Continuous Integrate-and-Fire (CIF) mechanism provides effective alignment for non-autoregressive (NAR) speech recognition. This mechanism creates a smooth and monotonic mapping from acoustic features to target tokens, achieving performance on Mandarin competitive with other NAR approaches. However, without finer-grained guidance, its stability degrades in some languages such as English and French. In this paper, we propose Multi-scale CIF (M-CIF), which performs multi-level alignment by integrating character and phoneme level supervision progressively distilled into subword representations, thereby enhancing robust acoustic-text alignment. Experiments show that M-CIF reduces WER compared to the Paraformer baseline, especially on CommonVoice by 4.21% in German and 3.05% in French. To further investigate these gains, we define phonetic confusion errors (PE) and space-related segmentation errors (SE) as evaluation metrics. Analysis of these metrics across different M-CIF settings reveals that the phoneme and character layers are essential for enhancing progressive CIF alignment. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_22172 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | M-CIF: Multi-Scale Alignment For CIF-Based Non-Autoregressive ASR Mao, Ruixiang Ma, Xiangnan Yang, Qing Zhu, Ziming Qiao, Yucheng Ge, Yuan Xiao, Tong Gao, Shengxiang Yu, Zhengtao Zhu, Jingbo Sound Computation and Language The Continuous Integrate-and-Fire (CIF) mechanism provides effective alignment for non-autoregressive (NAR) speech recognition. This mechanism creates a smooth and monotonic mapping from acoustic features to target tokens, achieving performance on Mandarin competitive with other NAR approaches. However, without finer-grained guidance, its stability degrades in some languages such as English and French. In this paper, we propose Multi-scale CIF (M-CIF), which performs multi-level alignment by integrating character and phoneme level supervision progressively distilled into subword representations, thereby enhancing robust acoustic-text alignment. Experiments show that M-CIF reduces WER compared to the Paraformer baseline, especially on CommonVoice by 4.21% in German and 3.05% in French. To further investigate these gains, we define phonetic confusion errors (PE) and space-related segmentation errors (SE) as evaluation metrics. Analysis of these metrics across different M-CIF settings reveals that the phoneme and character layers are essential for enhancing progressive CIF alignment. |
| title | M-CIF: Multi-Scale Alignment For CIF-Based Non-Autoregressive ASR |
| topic | Sound Computation and Language |
| url | https://arxiv.org/abs/2510.22172 |