TurboBias: Universal ASR Context-Biasing powered by GPU-accelerated Phrase-Boosting Tree
Fuente:
arXiv
Salvato in:
| Autori principali: | Andrusenko, Andrei, Bataev, Vladimir, Grigoryan, Lilit, Lavrukhin, Vitaly, Ginsburg, Boris |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
NGPU-LM: GPU-Accelerated N-Gram Language Model for Context-Biasing in Greedy ASR Decoding
di: Bataev, Vladimir, et al.
Pubblicazione: (2025)
di: Bataev, Vladimir, et al.
Pubblicazione: (2025)
FlexCTC: GPU-powered CTC Beam Decoding With Advanced Contextual Abilities
di: Grigoryan, Lilit, et al.
Pubblicazione: (2025)
di: Grigoryan, Lilit, et al.
Pubblicazione: (2025)
Pushing the Limits of Beam Search Decoding for Transducer-based ASR models
di: Grigoryan, Lilit, et al.
Pubblicazione: (2025)
di: Grigoryan, Lilit, et al.
Pubblicazione: (2025)
Fast Context-Biasing for CTC and Transducer ASR models with CTC-based Word Spotter
di: Andrusenko, Andrei, et al.
Pubblicazione: (2024)
di: Andrusenko, Andrei, et al.
Pubblicazione: (2024)
Reducing the Offline-Streaming Gap for Unified ASR Transducer with Consistency Regularization
di: Andrusenko, Andrei, et al.
Pubblicazione: (2026)
di: Andrusenko, Andrei, et al.
Pubblicazione: (2026)
Text-only domain adaptation for end-to-end ASR using integrated text-to-mel-spectrogram generator
di: Bataev, Vladimir, et al.
Pubblicazione: (2023)
di: Bataev, Vladimir, et al.
Pubblicazione: (2023)
Label-Looping: Highly Efficient Decoding for Transducers
di: Bataev, Vladimir, et al.
Pubblicazione: (2024)
di: Bataev, Vladimir, et al.
Pubblicazione: (2024)
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
di: Bataev, Vladimir, et al.
Pubblicazione: (2025)
di: Bataev, Vladimir, et al.
Pubblicazione: (2025)
Open Automatic Speech Recognition Models for Classical and Modern Standard Arabic
di: Grigoryan, Lilit, et al.
Pubblicazione: (2025)
di: Grigoryan, Lilit, et al.
Pubblicazione: (2025)
Methods to Increase the Amount of Data for Speech Recognition for Low Resource Languages
di: Ayrapetyan, Alexan, et al.
Pubblicazione: (2025)
di: Ayrapetyan, Alexan, et al.
Pubblicazione: (2025)
BR-ASR: Efficient and Scalable Bias Retrieval Framework for Contextual Biasing ASR in Speech LLM
di: Gong, Xun, et al.
Pubblicazione: (2025)
di: Gong, Xun, et al.
Pubblicazione: (2025)
Romanization Encoding For Multilingual ASR
di: Ding, Wen, et al.
Pubblicazione: (2024)
di: Ding, Wen, et al.
Pubblicazione: (2024)
Resource-Efficient Adaptation of Speech Foundation Models for Multi-Speaker ASR
di: Wang, Weiqing, et al.
Pubblicazione: (2024)
di: Wang, Weiqing, et al.
Pubblicazione: (2024)
Speaker Targeting via Self-Speaker Adaptation for Multi-talker ASR
di: Wang, Weiqing, et al.
Pubblicazione: (2025)
di: Wang, Weiqing, et al.
Pubblicazione: (2025)
EMMeTT: Efficient Multimodal Machine Translation Training
di: Żelasko, Piotr, et al.
Pubblicazione: (2024)
di: Żelasko, Piotr, et al.
Pubblicazione: (2024)
Contextualized Automatic Speech Recognition with Attention-Based Bias Phrase Boosted Beam Search
di: Sudo, Yui, et al.
Pubblicazione: (2024)
di: Sudo, Yui, et al.
Pubblicazione: (2024)
META-CAT: Speaker-Informed Speech Embeddings via Meta Information Concatenation for Multi-talker ASR
di: Wang, Jinhan, et al.
Pubblicazione: (2024)
di: Wang, Jinhan, et al.
Pubblicazione: (2024)
RNN-Transducer-based Losses for Speech Recognition on Noisy Targets
di: Bataev, Vladimir
Pubblicazione: (2025)
di: Bataev, Vladimir
Pubblicazione: (2025)
Unified Semi-Supervised Pipeline for Automatic Speech Recognition
di: Tadevosyan, Nune, et al.
Pubblicazione: (2025)
di: Tadevosyan, Nune, et al.
Pubblicazione: (2025)
PHRASED: Phrase Dictionary Biasing for Speech Translation
di: Wang, Peidong, et al.
Pubblicazione: (2025)
di: Wang, Peidong, et al.
Pubblicazione: (2025)
WCTC-Biasing: Retraining-free Contextual Biasing ASR with Wildcard CTC-based Keyword Spotting and Inter-layer Biasing
di: Nakagome, Yu, et al.
Pubblicazione: (2025)
di: Nakagome, Yu, et al.
Pubblicazione: (2025)
Right Label Context in End-to-End Training of Time-Synchronous ASR Models
di: Raissi, Tina, et al.
Pubblicazione: (2025)
di: Raissi, Tina, et al.
Pubblicazione: (2025)
Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
di: Bai, Ye, et al.
Pubblicazione: (2024)
di: Bai, Ye, et al.
Pubblicazione: (2024)
An approach to measuring the performance of Automatic Speech Recognition (ASR) models in the context of Large Language Model (LLM) powered applications
di: Pulikodan, Sujith, et al.
Pubblicazione: (2025)
di: Pulikodan, Sujith, et al.
Pubblicazione: (2025)
Minimising Biasing Word Errors for Contextual ASR with the Tree-Constrained Pointer Generator
di: Sun, Guangzhi, et al.
Pubblicazione: (2022)
di: Sun, Guangzhi, et al.
Pubblicazione: (2022)
Target Speaker ASR with Whisper
di: Polok, Alexander, et al.
Pubblicazione: (2024)
di: Polok, Alexander, et al.
Pubblicazione: (2024)
Index-ASR Technical Report
di: Song, Zheshu, et al.
Pubblicazione: (2025)
di: Song, Zheshu, et al.
Pubblicazione: (2025)
ASR Under the Stethoscope: Evaluating Biases in Clinical Speech Recognition across Indian Languages
di: Kumar, Subham, et al.
Pubblicazione: (2025)
di: Kumar, Subham, et al.
Pubblicazione: (2025)
Semi-Autoregressive Streaming ASR With Label Context
di: Arora, Siddhant, et al.
Pubblicazione: (2023)
di: Arora, Siddhant, et al.
Pubblicazione: (2023)
Speech Emotion Recognition with ASR Integration
di: Li, Yuanchao
Pubblicazione: (2026)
di: Li, Yuanchao
Pubblicazione: (2026)
Efficient Scaling for LLM-based ASR
di: Mu, Bingshen, et al.
Pubblicazione: (2025)
di: Mu, Bingshen, et al.
Pubblicazione: (2025)
String Sound Synthesizer on GPU-accelerated Finite Difference Scheme
di: Lee, Jin Woo, et al.
Pubblicazione: (2023)
di: Lee, Jin Woo, et al.
Pubblicazione: (2023)
Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering
di: Medennikov, Ivan, et al.
Pubblicazione: (2025)
di: Medennikov, Ivan, et al.
Pubblicazione: (2025)
Less is More: Accurate Speech Recognition & Translation without Web-Scale Data
di: Puvvada, Krishna C., et al.
Pubblicazione: (2024)
di: Puvvada, Krishna C., et al.
Pubblicazione: (2024)
LCB-net: Long-Context Biasing for Audio-Visual Speech Recognition
di: Yu, Fan, et al.
Pubblicazione: (2024)
di: Yu, Fan, et al.
Pubblicazione: (2024)
The USTC-NERCSLIP Systems for The ICMC-ASR Challenge
di: Wu, Minghui, et al.
Pubblicazione: (2024)
di: Wu, Minghui, et al.
Pubblicazione: (2024)
CatchPhrase: EXPrompt-Guided Encoder Adaptation for Audio-to-Image Generation
di: Oh, Hyunwoo, et al.
Pubblicazione: (2025)
di: Oh, Hyunwoo, et al.
Pubblicazione: (2025)
LUPET: Incorporating Hierarchical Information Path into Multilingual ASR
di: Liu, Wei, et al.
Pubblicazione: (2024)
di: Liu, Wei, et al.
Pubblicazione: (2024)
persoDA: Personalized Data Augmentation for Personalized ASR
di: Parada, Pablo Peso, et al.
Pubblicazione: (2025)
di: Parada, Pablo Peso, et al.
Pubblicazione: (2025)
Speaker Adaptation for Quantised End-to-End ASR Models
di: Zhao, Qiuming, et al.
Pubblicazione: (2024)
di: Zhao, Qiuming, et al.
Pubblicazione: (2024)
Documenti analoghi
-
NGPU-LM: GPU-Accelerated N-Gram Language Model for Context-Biasing in Greedy ASR Decoding
di: Bataev, Vladimir, et al.
Pubblicazione: (2025) -
FlexCTC: GPU-powered CTC Beam Decoding With Advanced Contextual Abilities
di: Grigoryan, Lilit, et al.
Pubblicazione: (2025) -
Pushing the Limits of Beam Search Decoding for Transducer-based ASR models
di: Grigoryan, Lilit, et al.
Pubblicazione: (2025) -
Fast Context-Biasing for CTC and Transducer ASR models with CTC-based Word Spotter
di: Andrusenko, Andrei, et al.
Pubblicazione: (2024) -
Reducing the Offline-Streaming Gap for Unified ASR Transducer with Consistency Regularization
di: Andrusenko, Andrei, et al.
Pubblicazione: (2026)