FlexCTC: GPU-powered CTC Beam Decoding With Advanced Contextual Abilities
Fuente:
arXiv
Guardado en:
| Autores principales: | Grigoryan, Lilit, Bataev, Vladimir, Karpov, Nikolay, Andrusenko, Andrei, Lavrukhin, Vitaly, Ginsburg, Boris |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
TurboBias: Universal ASR Context-Biasing powered by GPU-accelerated Phrase-Boosting Tree
por: Andrusenko, Andrei, et al.
Publicado: (2025)
por: Andrusenko, Andrei, et al.
Publicado: (2025)
Fast Context-Biasing for CTC and Transducer ASR models with CTC-based Word Spotter
por: Andrusenko, Andrei, et al.
Publicado: (2024)
por: Andrusenko, Andrei, et al.
Publicado: (2024)
Pushing the Limits of Beam Search Decoding for Transducer-based ASR models
por: Grigoryan, Lilit, et al.
Publicado: (2025)
por: Grigoryan, Lilit, et al.
Publicado: (2025)
NGPU-LM: GPU-Accelerated N-Gram Language Model for Context-Biasing in Greedy ASR Decoding
por: Bataev, Vladimir, et al.
Publicado: (2025)
por: Bataev, Vladimir, et al.
Publicado: (2025)
Reducing the Offline-Streaming Gap for Unified ASR Transducer with Consistency Regularization
por: Andrusenko, Andrei, et al.
Publicado: (2026)
por: Andrusenko, Andrei, et al.
Publicado: (2026)
Label-Looping: Highly Efficient Decoding for Transducers
por: Bataev, Vladimir, et al.
Publicado: (2024)
por: Bataev, Vladimir, et al.
Publicado: (2024)
Open Automatic Speech Recognition Models for Classical and Modern Standard Arabic
por: Grigoryan, Lilit, et al.
Publicado: (2025)
por: Grigoryan, Lilit, et al.
Publicado: (2025)
Text-only domain adaptation for end-to-end ASR using integrated text-to-mel-spectrogram generator
por: Bataev, Vladimir, et al.
Publicado: (2023)
por: Bataev, Vladimir, et al.
Publicado: (2023)
Joint Beam Search Integrating CTC, Attention, and Transducer Decoders
por: Sudo, Yui, et al.
Publicado: (2024)
por: Sudo, Yui, et al.
Publicado: (2024)
kNN-CTC: Enhancing ASR via Retrieval of CTC Pseudo Labels
por: Zhou, Jiaming, et al.
Publicado: (2023)
por: Zhou, Jiaming, et al.
Publicado: (2023)
CTC Blank Triggered Dynamic Layer-Skipping for Efficient CTC-based Speech Recognition
por: Hou, Junfeng, et al.
Publicado: (2024)
por: Hou, Junfeng, et al.
Publicado: (2024)
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
por: Bataev, Vladimir, et al.
Publicado: (2025)
por: Bataev, Vladimir, et al.
Publicado: (2025)
Methods to Increase the Amount of Data for Speech Recognition for Low Resource Languages
por: Ayrapetyan, Alexan, et al.
Publicado: (2025)
por: Ayrapetyan, Alexan, et al.
Publicado: (2025)
CR-CTC: Consistency regularization on CTC for improved speech recognition
por: Yao, Zengwei, et al.
Publicado: (2024)
por: Yao, Zengwei, et al.
Publicado: (2024)
Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation
por: Tsunoo, Emiru, et al.
Publicado: (2023)
por: Tsunoo, Emiru, et al.
Publicado: (2023)
CTC-TTS: LLM-based dual-streaming text-to-speech with CTC alignment
por: Liu, Hanwen, et al.
Publicado: (2026)
por: Liu, Hanwen, et al.
Publicado: (2026)
Unified Semi-Supervised Pipeline for Automatic Speech Recognition
por: Tadevosyan, Nune, et al.
Publicado: (2025)
por: Tadevosyan, Nune, et al.
Publicado: (2025)
Speaker-Distinguishable CTC: Learning Speaker Distinction Using CTC for Multi-Talker Speech Recognition
por: Sakuma, Asahi, et al.
Publicado: (2025)
por: Sakuma, Asahi, et al.
Publicado: (2025)
NTC-KWS: Noise-aware CTC for Robust Keyword Spotting
por: Xi, Yu, et al.
Publicado: (2024)
por: Xi, Yu, et al.
Publicado: (2024)
CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR
por: Zhou, Wei, et al.
Publicado: (2024)
por: Zhou, Wei, et al.
Publicado: (2024)
Unimodal Aggregation for CTC-based Speech Recognition
por: Fang, Ying, et al.
Publicado: (2023)
por: Fang, Ying, et al.
Publicado: (2023)
Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency Streaming ASR
por: Li, Longhao, et al.
Publicado: (2025)
por: Li, Longhao, et al.
Publicado: (2025)
Multilingual Audio-Visual Speech Recognition with Hybrid CTC/RNN-T Fast Conformer
por: Burchi, Maxime, et al.
Publicado: (2024)
por: Burchi, Maxime, et al.
Publicado: (2024)
Analyzing the Importance of Blank for CTC-Based Knowledge Distillation
por: Hilmes, Benedikt, et al.
Publicado: (2025)
por: Hilmes, Benedikt, et al.
Publicado: (2025)
SegAug: CTC-Aligned Segmented Augmentation For Robust RNN-Transducer Based Speech Recognition
por: Le, Khanh, et al.
Publicado: (2025)
por: Le, Khanh, et al.
Publicado: (2025)
Enhancing CTC-based speech recognition with diverse modeling units
por: Han, Shiyi, et al.
Publicado: (2024)
por: Han, Shiyi, et al.
Publicado: (2024)
WCTC-Biasing: Retraining-free Contextual Biasing ASR with Wildcard CTC-based Keyword Spotting and Inter-layer Biasing
por: Nakagome, Yu, et al.
Publicado: (2025)
por: Nakagome, Yu, et al.
Publicado: (2025)
LegoSLM: Connecting LLM with Speech Encoder using CTC Posteriors
por: Ma, Rao, et al.
Publicado: (2025)
por: Ma, Rao, et al.
Publicado: (2025)
Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
por: Kang, Jiawen, et al.
Publicado: (2024)
por: Kang, Jiawen, et al.
Publicado: (2024)
Enhancing CTC-Based Visual Speech Recognition
por: Laux, Hendrik, et al.
Publicado: (2024)
por: Laux, Hendrik, et al.
Publicado: (2024)
Boosting CTC-Based ASR Using LLM-Based Intermediate Loss Regularization
por: Altinok, Duygu
Publicado: (2025)
por: Altinok, Duygu
Publicado: (2025)
CTC-aligned Audio-Text Embedding for Streaming Open-vocabulary Keyword Spotting
por: Jin, Sichen, et al.
Publicado: (2024)
por: Jin, Sichen, et al.
Publicado: (2024)
FLToP CTC: Frame-Level Token Pruning via Relative Threshold for Efficient and Memory-Saving Decoding on Diverse Platforms
por: Shree, Atul, et al.
Publicado: (2025)
por: Shree, Atul, et al.
Publicado: (2025)
Label-Context-Dependent Internal Language Model Estimation for CTC
por: Yang, Zijian, et al.
Publicado: (2025)
por: Yang, Zijian, et al.
Publicado: (2025)
A Language-Agnostic Hierarchical LoRA-MoE Architecture for CTC-based Multilingual ASR
por: Zheng, Yuang, et al.
Publicado: (2026)
por: Zheng, Yuang, et al.
Publicado: (2026)
OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
por: Peng, Yifan, et al.
Publicado: (2024)
por: Peng, Yifan, et al.
Publicado: (2024)
EMMeTT: Efficient Multimodal Machine Translation Training
por: Żelasko, Piotr, et al.
Publicado: (2024)
por: Żelasko, Piotr, et al.
Publicado: (2024)
Guiding Frame-Level CTC Alignments Using Self-knowledge Distillation
por: Kim, Eungbeom, et al.
Publicado: (2024)
por: Kim, Eungbeom, et al.
Publicado: (2024)
Improving Zero-Shot Chinese-English Code-Switching ASR with kNN-CTC and Gated Monolingual Datastores
por: Zhou, Jiaming, et al.
Publicado: (2024)
por: Zhou, Jiaming, et al.
Publicado: (2024)
Improving Multilingual Speech Models on ML-SUPERB 2.0: Fine-tuning with Data Augmentation and LID-Aware CTC
por: Wang, Qingzheng, et al.
Publicado: (2025)
por: Wang, Qingzheng, et al.
Publicado: (2025)
Ejemplares similares
-
TurboBias: Universal ASR Context-Biasing powered by GPU-accelerated Phrase-Boosting Tree
por: Andrusenko, Andrei, et al.
Publicado: (2025) -
Fast Context-Biasing for CTC and Transducer ASR models with CTC-based Word Spotter
por: Andrusenko, Andrei, et al.
Publicado: (2024) -
Pushing the Limits of Beam Search Decoding for Transducer-based ASR models
por: Grigoryan, Lilit, et al.
Publicado: (2025) -
NGPU-LM: GPU-Accelerated N-Gram Language Model for Context-Biasing in Greedy ASR Decoding
por: Bataev, Vladimir, et al.
Publicado: (2025) -
Reducing the Offline-Streaming Gap for Unified ASR Transducer with Consistency Regularization
por: Andrusenko, Andrei, et al.
Publicado: (2026)