WIND: Accelerated RNN-T Decoding with Windowed Inference for Non-blank Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Hainan, Bataev, Vladimir, Grigoryan, Lilit, Ginsburg, Boris |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pushing the Limits of Beam Search Decoding for Transducer-based ASR models
by: Grigoryan, Lilit, et al.
Published: (2025)
by: Grigoryan, Lilit, et al.
Published: (2025)
NGPU-LM: GPU-Accelerated N-Gram Language Model for Context-Biasing in Greedy ASR Decoding
by: Bataev, Vladimir, et al.
Published: (2025)
by: Bataev, Vladimir, et al.
Published: (2025)
FlexCTC: GPU-powered CTC Beam Decoding With Advanced Contextual Abilities
by: Grigoryan, Lilit, et al.
Published: (2025)
by: Grigoryan, Lilit, et al.
Published: (2025)
Speed of Light Exact Greedy Decoding for RNN-T Speech Recognition Models on GPU
by: Galvez, Daniel, et al.
Published: (2024)
by: Galvez, Daniel, et al.
Published: (2024)
Label-Looping: Highly Efficient Decoding for Transducers
by: Bataev, Vladimir, et al.
Published: (2024)
by: Bataev, Vladimir, et al.
Published: (2024)
HAINAN: Fast and Accurate Transducer for Hybrid-Autoregressive ASR
by: Xu, Hainan, et al.
Published: (2024)
by: Xu, Hainan, et al.
Published: (2024)
Multi-blank Transducers for Speech Recognition
by: Xu, Hainan, et al.
Published: (2022)
by: Xu, Hainan, et al.
Published: (2022)
TurboBias: Universal ASR Context-Biasing powered by GPU-accelerated Phrase-Boosting Tree
by: Andrusenko, Andrei, et al.
Published: (2025)
by: Andrusenko, Andrei, et al.
Published: (2025)
RNN-Transducer-based Losses for Speech Recognition on Noisy Targets
by: Bataev, Vladimir
Published: (2025)
by: Bataev, Vladimir
Published: (2025)
Reducing the Offline-Streaming Gap for Unified ASR Transducer with Consistency Regularization
by: Andrusenko, Andrei, et al.
Published: (2026)
by: Andrusenko, Andrei, et al.
Published: (2026)
Chunk-wise Attention Transducers for Fast and Accurate Streaming Speech-to-Text
by: Xu, Hainan, et al.
Published: (2026)
by: Xu, Hainan, et al.
Published: (2026)
Text-only domain adaptation for end-to-end ASR using integrated text-to-mel-spectrogram generator
by: Bataev, Vladimir, et al.
Published: (2023)
by: Bataev, Vladimir, et al.
Published: (2023)
Fast Context-Biasing for CTC and Transducer ASR models with CTC-based Word Spotter
by: Andrusenko, Andrei, et al.
Published: (2024)
by: Andrusenko, Andrei, et al.
Published: (2024)
Transducers with Pronunciation-aware Embeddings for Automatic Speech Recognition
by: Xu, Hainan, et al.
Published: (2024)
by: Xu, Hainan, et al.
Published: (2024)
Open Automatic Speech Recognition Models for Classical and Modern Standard Arabic
by: Grigoryan, Lilit, et al.
Published: (2025)
by: Grigoryan, Lilit, et al.
Published: (2025)
Star Attention: Efficient LLM Inference over Long Sequences
by: Acharya, Shantanu, et al.
Published: (2024)
by: Acharya, Shantanu, et al.
Published: (2024)
Window-Diffusion: Accelerating Diffusion Language Model Inference with Windowed Token Pruning and Caching
by: Zuo, Fengrui, et al.
Published: (2026)
by: Zuo, Fengrui, et al.
Published: (2026)
Activity Sparsity Complements Weight Sparsity for Efficient RNN Inference
by: Mukherji, Rishav, et al.
Published: (2023)
by: Mukherji, Rishav, et al.
Published: (2023)
Faster WIND: Accelerating Iterative Best-of-$N$ Distillation for LLM Alignment
by: Yang, Tong, et al.
Published: (2024)
by: Yang, Tong, et al.
Published: (2024)
Attention as an RNN
by: Feng, Leo, et al.
Published: (2024)
by: Feng, Leo, et al.
Published: (2024)
Set Block Decoding is a Language Model Inference Accelerator
by: Gat, Itai, et al.
Published: (2025)
by: Gat, Itai, et al.
Published: (2025)
SpecPipe: Accelerating Pipeline Parallelism-based LLM Inference with Speculative Decoding
by: Yin, Haofei, et al.
Published: (2025)
by: Yin, Haofei, et al.
Published: (2025)
SKOLR: Structured Koopman Operator Linear RNN for Time-Series Forecasting
by: Zhang, Yitian, et al.
Published: (2025)
by: Zhang, Yitian, et al.
Published: (2025)
Accelerating Transformer Inference for Translation via Parallel Decoding
by: Santilli, Andrea, et al.
Published: (2023)
by: Santilli, Andrea, et al.
Published: (2023)
Accelerating Inference of Discrete Autoregressive Normalizing Flows by Selective Jacobi Decoding
by: Zhang, Jiaru, et al.
Published: (2025)
by: Zhang, Jiaru, et al.
Published: (2025)
Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
by: Cai, Tianle, et al.
Published: (2024)
by: Cai, Tianle, et al.
Published: (2024)
CAS-Spec: Cascade Adaptive Self-Speculative Decoding for On-the-Fly Lossless Inference Acceleration of LLMs
by: Ning, Zhiyuan, et al.
Published: (2025)
by: Ning, Zhiyuan, et al.
Published: (2025)
Neural Window Decoder for SC-LDPC Codes
by: Yun, Dae-Young, et al.
Published: (2024)
by: Yun, Dae-Young, et al.
Published: (2024)
POSEIDON: Physics-Optimized Seismic Energy Inference and Detection Operating Network
by: Kriuk, Boris, et al.
Published: (2026)
by: Kriuk, Boris, et al.
Published: (2026)
RotRNN: Modelling Long Sequences with Rotations
by: Biegun, Kai, et al.
Published: (2024)
by: Biegun, Kai, et al.
Published: (2024)
Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding
by: Zhao, Yilong, et al.
Published: (2025)
by: Zhao, Yilong, et al.
Published: (2025)
Recursive Speculative Decoding: Accelerating LLM Inference via Sampling Without Replacement
by: Jeon, Wonseok, et al.
Published: (2024)
by: Jeon, Wonseok, et al.
Published: (2024)
Hardware-Aware Parallel Prompt Decoding for Memory-Efficient Acceleration of LLM Inference
by: Chen, Hao Mark, et al.
Published: (2024)
by: Chen, Hao Mark, et al.
Published: (2024)
Bayesian Inference with Shaped Deep Non-linear MLPs
by: Hanin, Boris, et al.
Published: (2026)
by: Hanin, Boris, et al.
Published: (2026)
ExARNN: An Environment-Driven Adaptive RNN for Learning Non-Stationary Power Dynamics
by: Li, Haoran, et al.
Published: (2025)
by: Li, Haoran, et al.
Published: (2025)
Attention Augmented GNN RNN-Attention Models for Advanced Cybersecurity Intrusion Detection
by: Biradar, Jayant, et al.
Published: (2025)
by: Biradar, Jayant, et al.
Published: (2025)
Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies
by: Timor, Nadav, et al.
Published: (2025)
by: Timor, Nadav, et al.
Published: (2025)
Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration
by: Wen, Zhuofan, et al.
Published: (2024)
by: Wen, Zhuofan, et al.
Published: (2024)
M$^2$RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling
by: Mishra, Mayank, et al.
Published: (2026)
by: Mishra, Mayank, et al.
Published: (2026)
Tighter Truncated Rectangular Prism Approximation for RNN Robustness Verification
by: Lin, Xingqi, et al.
Published: (2025)
by: Lin, Xingqi, et al.
Published: (2025)
Similar Items
-
Pushing the Limits of Beam Search Decoding for Transducer-based ASR models
by: Grigoryan, Lilit, et al.
Published: (2025) -
NGPU-LM: GPU-Accelerated N-Gram Language Model for Context-Biasing in Greedy ASR Decoding
by: Bataev, Vladimir, et al.
Published: (2025) -
FlexCTC: GPU-powered CTC Beam Decoding With Advanced Contextual Abilities
by: Grigoryan, Lilit, et al.
Published: (2025) -
Speed of Light Exact Greedy Decoding for RNN-T Speech Recognition Models on GPU
by: Galvez, Daniel, et al.
Published: (2024) -
Label-Looping: Highly Efficient Decoding for Transducers
by: Bataev, Vladimir, et al.
Published: (2024)