Label-Looping: Highly Efficient Decoding for Transducers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bataev, Vladimir, Xu, Hainan, Galvez, Daniel, Lavrukhin, Vitaly, Ginsburg, Boris |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Pushing the Limits of Beam Search Decoding for Transducer-based ASR models
von: Grigoryan, Lilit, et al.
Veröffentlicht: (2025)
von: Grigoryan, Lilit, et al.
Veröffentlicht: (2025)
Fast Context-Biasing for CTC and Transducer ASR models with CTC-based Word Spotter
von: Andrusenko, Andrei, et al.
Veröffentlicht: (2024)
von: Andrusenko, Andrei, et al.
Veröffentlicht: (2024)
FlexCTC: GPU-powered CTC Beam Decoding With Advanced Contextual Abilities
von: Grigoryan, Lilit, et al.
Veröffentlicht: (2025)
von: Grigoryan, Lilit, et al.
Veröffentlicht: (2025)
NGPU-LM: GPU-Accelerated N-Gram Language Model for Context-Biasing in Greedy ASR Decoding
von: Bataev, Vladimir, et al.
Veröffentlicht: (2025)
von: Bataev, Vladimir, et al.
Veröffentlicht: (2025)
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
von: Bataev, Vladimir, et al.
Veröffentlicht: (2025)
von: Bataev, Vladimir, et al.
Veröffentlicht: (2025)
TurboBias: Universal ASR Context-Biasing powered by GPU-accelerated Phrase-Boosting Tree
von: Andrusenko, Andrei, et al.
Veröffentlicht: (2025)
von: Andrusenko, Andrei, et al.
Veröffentlicht: (2025)
Text-only domain adaptation for end-to-end ASR using integrated text-to-mel-spectrogram generator
von: Bataev, Vladimir, et al.
Veröffentlicht: (2023)
von: Bataev, Vladimir, et al.
Veröffentlicht: (2023)
RNN-Transducer-based Losses for Speech Recognition on Noisy Targets
von: Bataev, Vladimir
Veröffentlicht: (2025)
von: Bataev, Vladimir
Veröffentlicht: (2025)
Transducers with Pronunciation-aware Embeddings for Automatic Speech Recognition
von: Xu, Hainan, et al.
Veröffentlicht: (2024)
von: Xu, Hainan, et al.
Veröffentlicht: (2024)
Reducing the Offline-Streaming Gap for Unified ASR Transducer with Consistency Regularization
von: Andrusenko, Andrei, et al.
Veröffentlicht: (2026)
von: Andrusenko, Andrei, et al.
Veröffentlicht: (2026)
Multi-blank Transducers for Speech Recognition
von: Xu, Hainan, et al.
Veröffentlicht: (2022)
von: Xu, Hainan, et al.
Veröffentlicht: (2022)
EMMeTT: Efficient Multimodal Machine Translation Training
von: Żelasko, Piotr, et al.
Veröffentlicht: (2024)
von: Żelasko, Piotr, et al.
Veröffentlicht: (2024)
Less is More: Accurate Speech Recognition & Translation without Web-Scale Data
von: Puvvada, Krishna C., et al.
Veröffentlicht: (2024)
von: Puvvada, Krishna C., et al.
Veröffentlicht: (2024)
Methods to Increase the Amount of Data for Speech Recognition for Low Resource Languages
von: Ayrapetyan, Alexan, et al.
Veröffentlicht: (2025)
von: Ayrapetyan, Alexan, et al.
Veröffentlicht: (2025)
Aligner-Encoders: Self-Attention Transformers Can Be Self-Transducers
von: Stooke, Adam, et al.
Veröffentlicht: (2025)
von: Stooke, Adam, et al.
Veröffentlicht: (2025)
Romanization Encoding For Multilingual ASR
von: Ding, Wen, et al.
Veröffentlicht: (2024)
von: Ding, Wen, et al.
Veröffentlicht: (2024)
Whisper in Medusa's Ear: Multi-head Efficient Decoding for Transformer-based ASR
von: Segal-Feldman, Yael, et al.
Veröffentlicht: (2024)
von: Segal-Feldman, Yael, et al.
Veröffentlicht: (2024)
Decoding Poultry Vocalizations -- Natural Language Processing and Transformer Models for Semantic and Emotional Analysis
von: Manikandan, Venkatraman, et al.
Veröffentlicht: (2024)
von: Manikandan, Venkatraman, et al.
Veröffentlicht: (2024)
Large Language Model Based Generative Error Correction: A Challenge and Baselines for Speech Recognition, Speaker Tagging, and Emotion Recognition
von: Yang, Chao-Han Huck, et al.
Veröffentlicht: (2024)
von: Yang, Chao-Han Huck, et al.
Veröffentlicht: (2024)
FlashSpeech: Efficient Zero-Shot Speech Synthesis
von: Ye, Zhen, et al.
Veröffentlicht: (2024)
von: Ye, Zhen, et al.
Veröffentlicht: (2024)
MoonCast: High-Quality Zero-Shot Podcast Generation
von: Ju, Zeqian, et al.
Veröffentlicht: (2025)
von: Ju, Zeqian, et al.
Veröffentlicht: (2025)
Dual Knowledge Distillation for Efficient Sound Event Detection
von: Xiao, Yang, et al.
Veröffentlicht: (2024)
von: Xiao, Yang, et al.
Veröffentlicht: (2024)
EAT: Self-Supervised Pre-Training with Efficient Audio Transformer
von: Chen, Wenxi, et al.
Veröffentlicht: (2024)
von: Chen, Wenxi, et al.
Veröffentlicht: (2024)
Speculative End-Turn Detector for Efficient Speech Chatbot Assistant
von: Ok, Hyunjong, et al.
Veröffentlicht: (2025)
von: Ok, Hyunjong, et al.
Veröffentlicht: (2025)
QUADS: QUAntized Distillation Framework for Efficient Speech Language Understanding
von: Biswas, Subrata, et al.
Veröffentlicht: (2025)
von: Biswas, Subrata, et al.
Veröffentlicht: (2025)
Large Language Models are Efficient Learners of Noise-Robust Speech Recognition
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
Towards Human-in-the-Loop Onset Detection: A Transfer Learning Approach for Maracatu
von: Pinto, António Sá
Veröffentlicht: (2025)
von: Pinto, António Sá
Veröffentlicht: (2025)
Whisfusion: Parallel ASR Decoding via a Diffusion Transformer
von: Kwon, Taeyoun, et al.
Veröffentlicht: (2025)
von: Kwon, Taeyoun, et al.
Veröffentlicht: (2025)
ELP-Adapters: Parameter Efficient Adapter Tuning for Various Speech Processing Tasks
von: Inoue, Nakamasa, et al.
Veröffentlicht: (2024)
von: Inoue, Nakamasa, et al.
Veröffentlicht: (2024)
Dynamic ASR Pathways: An Adaptive Masking Approach Towards Efficient Pruning of A Multilingual ASR Model
von: Xie, Jiamin, et al.
Veröffentlicht: (2023)
von: Xie, Jiamin, et al.
Veröffentlicht: (2023)
Can a Machine Distinguish High and Low Amount of Social Creak in Speech?
von: Laukkanen, Anne-Maria, et al.
Veröffentlicht: (2024)
von: Laukkanen, Anne-Maria, et al.
Veröffentlicht: (2024)
Music2Latent2: Audio Compression with Summary Embeddings and Autoregressive Decoding
von: Pasini, Marco, et al.
Veröffentlicht: (2025)
von: Pasini, Marco, et al.
Veröffentlicht: (2025)
Streaming Speaker Change Detection and Gender Classification for Transducer-Based Multi-Talker Speech Translation
von: Wang, Peidong, et al.
Veröffentlicht: (2025)
von: Wang, Peidong, et al.
Veröffentlicht: (2025)
Speak, Edit, Repeat: High-Fidelity Voice Editing and Zero-Shot TTS with Cross-Attentive Mamba
von: Mohammad, Baher, et al.
Veröffentlicht: (2025)
von: Mohammad, Baher, et al.
Veröffentlicht: (2025)
Efficient and Microphone-Fault-Tolerant 3D Sound Source Localization
von: Yang, Yiyuan, et al.
Veröffentlicht: (2025)
von: Yang, Yiyuan, et al.
Veröffentlicht: (2025)
On the Relation between Internal Language Model and Sequence Discriminative Training for Neural Transducers
von: Yang, Zijian, et al.
Veröffentlicht: (2023)
von: Yang, Zijian, et al.
Veröffentlicht: (2023)
Utilizing Neural Transducers for Two-Stage Text-to-Speech via Semantic Token Prediction
von: Kim, Minchan, et al.
Veröffentlicht: (2024)
von: Kim, Minchan, et al.
Veröffentlicht: (2024)
Towards End-to-End Training of Automatic Speech Recognition for Nigerian Pidgin
von: Rufai, Amina Mardiyyah, et al.
Veröffentlicht: (2020)
von: Rufai, Amina Mardiyyah, et al.
Veröffentlicht: (2020)
TinyMusician: On-Device Music Generation with Knowledge Distillation and Mixed Precision Quantization
von: Wang, Hainan, et al.
Veröffentlicht: (2025)
von: Wang, Hainan, et al.
Veröffentlicht: (2025)
CSyMR: Benchmarking Compositional Music Information Retrieval in Symbolic Music Reasoning
von: Wang, Boyang, et al.
Veröffentlicht: (2025)
von: Wang, Boyang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Pushing the Limits of Beam Search Decoding for Transducer-based ASR models
von: Grigoryan, Lilit, et al.
Veröffentlicht: (2025) -
Fast Context-Biasing for CTC and Transducer ASR models with CTC-based Word Spotter
von: Andrusenko, Andrei, et al.
Veröffentlicht: (2024) -
FlexCTC: GPU-powered CTC Beam Decoding With Advanced Contextual Abilities
von: Grigoryan, Lilit, et al.
Veröffentlicht: (2025) -
NGPU-LM: GPU-Accelerated N-Gram Language Model for Context-Biasing in Greedy ASR Decoding
von: Bataev, Vladimir, et al.
Veröffentlicht: (2025) -
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
von: Bataev, Vladimir, et al.
Veröffentlicht: (2025)