FlexCTC: GPU-powered CTC Beam Decoding With Advanced Contextual Abilities
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916895296323584 |
|---|---|
| author | Grigoryan, Lilit Bataev, Vladimir Karpov, Nikolay Andrusenko, Andrei Lavrukhin, Vitaly Ginsburg, Boris |
| author_facet | Grigoryan, Lilit Bataev, Vladimir Karpov, Nikolay Andrusenko, Andrei Lavrukhin, Vitaly Ginsburg, Boris |
| contents | While beam search improves speech recognition quality over greedy decoding, standard implementations are slow, often sequential, and CPU-bound. To fully leverage modern hardware capabilities, we present a novel open-source FlexCTC toolkit for fully GPU-based beam decoding, designed for Connectionist Temporal Classification (CTC) models. Developed entirely in Python and PyTorch, it offers a fast, user-friendly, and extensible alternative to traditional C++, CUDA, or WFST-based decoders. The toolkit features a high-performance, fully batched GPU implementation with eliminated CPU-GPU synchronization and minimized kernel launch overhead via CUDA Graphs. It also supports advanced contextualization techniques, including GPU-powered N-gram language model fusion and phrase-level boosting. These features enable accurate and efficient decoding, making them suitable for both research and production use. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_07315 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | FlexCTC: GPU-powered CTC Beam Decoding With Advanced Contextual Abilities Grigoryan, Lilit Bataev, Vladimir Karpov, Nikolay Andrusenko, Andrei Lavrukhin, Vitaly Ginsburg, Boris Audio and Speech Processing Artificial Intelligence Computation and Language Machine Learning Sound While beam search improves speech recognition quality over greedy decoding, standard implementations are slow, often sequential, and CPU-bound. To fully leverage modern hardware capabilities, we present a novel open-source FlexCTC toolkit for fully GPU-based beam decoding, designed for Connectionist Temporal Classification (CTC) models. Developed entirely in Python and PyTorch, it offers a fast, user-friendly, and extensible alternative to traditional C++, CUDA, or WFST-based decoders. The toolkit features a high-performance, fully batched GPU implementation with eliminated CPU-GPU synchronization and minimized kernel launch overhead via CUDA Graphs. It also supports advanced contextualization techniques, including GPU-powered N-gram language model fusion and phrase-level boosting. These features enable accurate and efficient decoding, making them suitable for both research and production use. |
| title | FlexCTC: GPU-powered CTC Beam Decoding With Advanced Contextual Abilities |
| topic | Audio and Speech Processing Artificial Intelligence Computation and Language Machine Learning Sound |
| url | https://arxiv.org/abs/2508.07315 |