Chunk-wise Attention Transducers for Fast and Accurate Streaming Speech-to-Text
Fuente:
arXiv
Salvato in:
| Autori principali: | Xu, Hainan, Bataev, Vladimir, Bartley, Travis M., Balam, Jagadeesh |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
HAINAN: Fast and Accurate Transducer for Hybrid-Autoregressive ASR
di: Xu, Hainan, et al.
Pubblicazione: (2024)
di: Xu, Hainan, et al.
Pubblicazione: (2024)
RNN-Transducer-based Losses for Speech Recognition on Noisy Targets
di: Bataev, Vladimir
Pubblicazione: (2025)
di: Bataev, Vladimir
Pubblicazione: (2025)
Label-Looping: Highly Efficient Decoding for Transducers
di: Bataev, Vladimir, et al.
Pubblicazione: (2024)
di: Bataev, Vladimir, et al.
Pubblicazione: (2024)
Speed of Light Exact Greedy Decoding for RNN-T Speech Recognition Models on GPU
di: Galvez, Daniel, et al.
Pubblicazione: (2024)
di: Galvez, Daniel, et al.
Pubblicazione: (2024)
Pushing the Limits of Beam Search Decoding for Transducer-based ASR models
di: Grigoryan, Lilit, et al.
Pubblicazione: (2025)
di: Grigoryan, Lilit, et al.
Pubblicazione: (2025)
WIND: Accelerated RNN-T Decoding with Windowed Inference for Non-blank Detection
di: Xu, Hainan, et al.
Pubblicazione: (2025)
di: Xu, Hainan, et al.
Pubblicazione: (2025)
Multi-blank Transducers for Speech Recognition
di: Xu, Hainan, et al.
Pubblicazione: (2022)
di: Xu, Hainan, et al.
Pubblicazione: (2022)
Fast Context-Biasing for CTC and Transducer ASR models with CTC-based Word Spotter
di: Andrusenko, Andrei, et al.
Pubblicazione: (2024)
di: Andrusenko, Andrei, et al.
Pubblicazione: (2024)
Transducers with Pronunciation-aware Embeddings for Automatic Speech Recognition
di: Xu, Hainan, et al.
Pubblicazione: (2024)
di: Xu, Hainan, et al.
Pubblicazione: (2024)
Longer is (Not Necessarily) Stronger: Punctuated Long-Sequence Training for Enhanced Speech Recognition and Translation
di: Koluguri, Nithin Rao, et al.
Pubblicazione: (2024)
di: Koluguri, Nithin Rao, et al.
Pubblicazione: (2024)
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling
di: Sun, Haiyang, et al.
Pubblicazione: (2025)
di: Sun, Haiyang, et al.
Pubblicazione: (2025)
Chunked Attention-based Encoder-Decoder Model for Streaming Speech Recognition
di: Zeineldeen, Mohammad, et al.
Pubblicazione: (2023)
di: Zeineldeen, Mohammad, et al.
Pubblicazione: (2023)
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
di: Bataev, Vladimir, et al.
Pubblicazione: (2025)
di: Bataev, Vladimir, et al.
Pubblicazione: (2025)
Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition
di: Lee, Hyeonseung, et al.
Pubblicazione: (2024)
di: Lee, Hyeonseung, et al.
Pubblicazione: (2024)
Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations
di: Dhawan, Kunal, et al.
Pubblicazione: (2024)
di: Dhawan, Kunal, et al.
Pubblicazione: (2024)
Instruction Data Generation and Unsupervised Adaptation for Speech Language Models
di: Noroozi, Vahid, et al.
Pubblicazione: (2024)
di: Noroozi, Vahid, et al.
Pubblicazione: (2024)
Less is More: Accurate Speech Recognition & Translation without Web-Scale Data
di: Puvvada, Krishna C., et al.
Pubblicazione: (2024)
di: Puvvada, Krishna C., et al.
Pubblicazione: (2024)
Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems
di: Park, Taejin, et al.
Pubblicazione: (2024)
di: Park, Taejin, et al.
Pubblicazione: (2024)
Training Long-Context LLMs Efficiently via Chunk-wise Optimization
di: Li, Wenhao, et al.
Pubblicazione: (2025)
di: Li, Wenhao, et al.
Pubblicazione: (2025)
Text-only domain adaptation for end-to-end ASR using integrated text-to-mel-spectrogram generator
di: Bataev, Vladimir, et al.
Pubblicazione: (2023)
di: Bataev, Vladimir, et al.
Pubblicazione: (2023)
Delta Attention: Fast and Accurate Sparse Attention Inference by Delta Correction
di: Willette, Jeffrey, et al.
Pubblicazione: (2025)
di: Willette, Jeffrey, et al.
Pubblicazione: (2025)
CMoS: Rethinking Time Series Prediction Through the Lens of Chunk-wise Spatial Correlations
di: Si, Haotian, et al.
Pubblicazione: (2025)
di: Si, Haotian, et al.
Pubblicazione: (2025)
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
di: Shah, Jay, et al.
Pubblicazione: (2024)
di: Shah, Jay, et al.
Pubblicazione: (2024)
Fast and Accurate Triangle Counting in Graph Streams Using Predictions
di: Boldrin, Cristian, et al.
Pubblicazione: (2024)
di: Boldrin, Cristian, et al.
Pubblicazione: (2024)
Reducing the Offline-Streaming Gap for Unified ASR Transducer with Consistency Regularization
di: Andrusenko, Andrei, et al.
Pubblicazione: (2026)
di: Andrusenko, Andrei, et al.
Pubblicazione: (2026)
Utilizing Neural Transducers for Two-Stage Text-to-Speech via Semantic Token Prediction
di: Kim, Minchan, et al.
Pubblicazione: (2024)
di: Kim, Minchan, et al.
Pubblicazione: (2024)
SpeakStream: Streaming Text-to-Speech with Interleaved Data
di: Bai, Richard He, et al.
Pubblicazione: (2025)
di: Bai, Richard He, et al.
Pubblicazione: (2025)
Cognitive Chunking for Soft Prompts: Accelerating Compressor Learning via Block-wise Causal Masking
di: Liu, Guojie, et al.
Pubblicazione: (2026)
di: Liu, Guojie, et al.
Pubblicazione: (2026)
TDT-KWS: Fast And Accurate Keyword Spotting Using Token-and-duration Transducer
di: Xi, Yu, et al.
Pubblicazione: (2024)
di: Xi, Yu, et al.
Pubblicazione: (2024)
ChunkFT: Byte-Streamed Optimization for Memory-Efficient Full Fine-Tuning
di: Liu, Yongkang, et al.
Pubblicazione: (2026)
di: Liu, Yongkang, et al.
Pubblicazione: (2026)
Learning Self-Growth Maps for Fast and Accurate Imbalanced Streaming Data Clustering
di: Zhang, Yiqun, et al.
Pubblicazione: (2024)
di: Zhang, Yiqun, et al.
Pubblicazione: (2024)
Support Basis: Fast Attention Beyond Bounded Entries
di: Aliakbarpour, Maryam, et al.
Pubblicazione: (2025)
di: Aliakbarpour, Maryam, et al.
Pubblicazione: (2025)
ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition
di: Ye, Lu, et al.
Pubblicazione: (2024)
di: Ye, Lu, et al.
Pubblicazione: (2024)
SegINR: Segment-wise Implicit Neural Representation for Sequence Alignment in Neural Text-to-Speech
di: Kim, Minchan, et al.
Pubblicazione: (2024)
di: Kim, Minchan, et al.
Pubblicazione: (2024)
Fast Multipole Attention: A Scalable Multilevel Attention Mechanism for Text and Images
di: Kang, Yanming, et al.
Pubblicazione: (2023)
di: Kang, Yanming, et al.
Pubblicazione: (2023)
Provably Convergent Subgraph-wise Sampling for Fast GNN Training
di: Wang, Jie, et al.
Pubblicazione: (2023)
di: Wang, Jie, et al.
Pubblicazione: (2023)
CAFE-GB: Scalable and Stable Feature Selection for Malware Detection via Chunk-wise Aggregated Gradient Boosting
di: K, Ajvad Haneef, et al.
Pubblicazione: (2026)
di: K, Ajvad Haneef, et al.
Pubblicazione: (2026)
MeLeMaD: Adaptive Malware Detection via Chunk-wise Feature Selection and Meta-Learning
di: K, Ajvad Haneef, et al.
Pubblicazione: (2025)
di: K, Ajvad Haneef, et al.
Pubblicazione: (2025)
Cluster-wise Graph Transformer with Dual-granularity Kernelized Attention
di: Huang, Siyuan, et al.
Pubblicazione: (2024)
di: Huang, Siyuan, et al.
Pubblicazione: (2024)
Stateful Conformer with Cache-based Inference for Streaming Automatic Speech Recognition
di: Noroozi, Vahid, et al.
Pubblicazione: (2023)
di: Noroozi, Vahid, et al.
Pubblicazione: (2023)
Documenti analoghi
-
HAINAN: Fast and Accurate Transducer for Hybrid-Autoregressive ASR
di: Xu, Hainan, et al.
Pubblicazione: (2024) -
RNN-Transducer-based Losses for Speech Recognition on Noisy Targets
di: Bataev, Vladimir
Pubblicazione: (2025) -
Label-Looping: Highly Efficient Decoding for Transducers
di: Bataev, Vladimir, et al.
Pubblicazione: (2024) -
Speed of Light Exact Greedy Decoding for RNN-T Speech Recognition Models on GPU
di: Galvez, Daniel, et al.
Pubblicazione: (2024) -
Pushing the Limits of Beam Search Decoding for Transducer-based ASR models
di: Grigoryan, Lilit, et al.
Pubblicazione: (2025)