FlexCTC: GPU-powered CTC Beam Decoding With Advanced Contextual Abilities

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Grigoryan, Lilit, Bataev, Vladimir, Karpov, Nikolay, Andrusenko, Andrei, Lavrukhin, Vitaly, Ginsburg, Boris
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916895296323584
author Grigoryan, Lilit
Bataev, Vladimir
Karpov, Nikolay
Andrusenko, Andrei
Lavrukhin, Vitaly
Ginsburg, Boris
author_facet Grigoryan, Lilit
Bataev, Vladimir
Karpov, Nikolay
Andrusenko, Andrei
Lavrukhin, Vitaly
Ginsburg, Boris
contents While beam search improves speech recognition quality over greedy decoding, standard implementations are slow, often sequential, and CPU-bound. To fully leverage modern hardware capabilities, we present a novel open-source FlexCTC toolkit for fully GPU-based beam decoding, designed for Connectionist Temporal Classification (CTC) models. Developed entirely in Python and PyTorch, it offers a fast, user-friendly, and extensible alternative to traditional C++, CUDA, or WFST-based decoders. The toolkit features a high-performance, fully batched GPU implementation with eliminated CPU-GPU synchronization and minimized kernel launch overhead via CUDA Graphs. It also supports advanced contextualization techniques, including GPU-powered N-gram language model fusion and phrase-level boosting. These features enable accurate and efficient decoding, making them suitable for both research and production use.
format Preprint
id arxiv_https___arxiv_org_abs_2508_07315
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FlexCTC: GPU-powered CTC Beam Decoding With Advanced Contextual Abilities
Grigoryan, Lilit
Bataev, Vladimir
Karpov, Nikolay
Andrusenko, Andrei
Lavrukhin, Vitaly
Ginsburg, Boris
Audio and Speech Processing
Artificial Intelligence
Computation and Language
Machine Learning
Sound
While beam search improves speech recognition quality over greedy decoding, standard implementations are slow, often sequential, and CPU-bound. To fully leverage modern hardware capabilities, we present a novel open-source FlexCTC toolkit for fully GPU-based beam decoding, designed for Connectionist Temporal Classification (CTC) models. Developed entirely in Python and PyTorch, it offers a fast, user-friendly, and extensible alternative to traditional C++, CUDA, or WFST-based decoders. The toolkit features a high-performance, fully batched GPU implementation with eliminated CPU-GPU synchronization and minimized kernel launch overhead via CUDA Graphs. It also supports advanced contextualization techniques, including GPU-powered N-gram language model fusion and phrase-level boosting. These features enable accurate and efficient decoding, making them suitable for both research and production use.
title FlexCTC: GPU-powered CTC Beam Decoding With Advanced Contextual Abilities
topic Audio and Speech Processing
Artificial Intelligence
Computation and Language
Machine Learning
Sound
url https://arxiv.org/abs/2508.07315