CUSIDE-T: Chunking, Simulating Future and Decoding for Transducer based Streaming ASR
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Wenbo, Li, Ziwei, Yu, Chuan, Ou, Zhijian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Masked Self-distilled Transducer-based Keyword Spotting with Semi-autoregressive Decoding
von: Xi, Yu, et al.
Veröffentlicht: (2025)
von: Xi, Yu, et al.
Veröffentlicht: (2025)
SSCFormer: Push the Limit of Chunk-wise Conformer for Streaming ASR Using Sequentially Sampled Chunks and Chunked Causal Convolution
von: Wang, Fangyuan, et al.
Veröffentlicht: (2022)
von: Wang, Fangyuan, et al.
Veröffentlicht: (2022)
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
von: Thorbecke, Iuliia, et al.
Veröffentlicht: (2024)
von: Thorbecke, Iuliia, et al.
Veröffentlicht: (2024)
Chunked Attention-based Encoder-Decoder Model for Streaming Speech Recognition
von: Zeineldeen, Mohammad, et al.
Veröffentlicht: (2023)
von: Zeineldeen, Mohammad, et al.
Veröffentlicht: (2023)
VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech
von: Du, Chenpeng, et al.
Veröffentlicht: (2024)
von: Du, Chenpeng, et al.
Veröffentlicht: (2024)
CUSIDE-array: A Streaming Multi-Channel End-to-End Speech Recognition System with Realistic Evaluations
von: Kong, Xiangzhu, et al.
Veröffentlicht: (2024)
von: Kong, Xiangzhu, et al.
Veröffentlicht: (2024)
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
von: Dong, Lukuang, et al.
Veröffentlicht: (2026)
von: Dong, Lukuang, et al.
Veröffentlicht: (2026)
Promptformer: Prompted Conformer Transducer for ASR
von: Duarte-Torres, Sergio, et al.
Veröffentlicht: (2024)
von: Duarte-Torres, Sergio, et al.
Veröffentlicht: (2024)
Alignment-Free Training for Transducer-based Multi-Talker ASR
von: Moriya, Takafumi, et al.
Veröffentlicht: (2024)
von: Moriya, Takafumi, et al.
Veröffentlicht: (2024)
Phoneme-based speech recognition driven by large language models and sampling marginalization
von: Ma, Te, et al.
Veröffentlicht: (2025)
von: Ma, Te, et al.
Veröffentlicht: (2025)
Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency Streaming ASR
von: Li, Longhao, et al.
Veröffentlicht: (2025)
von: Li, Longhao, et al.
Veröffentlicht: (2025)
EfficientASR: Speech Recognition Network Compression via Attention Redundancy and Chunk-Level FFN Optimization
von: Wang, Jianzong, et al.
Veröffentlicht: (2024)
von: Wang, Jianzong, et al.
Veröffentlicht: (2024)
SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding
von: Wei, Linye, et al.
Veröffentlicht: (2025)
von: Wei, Linye, et al.
Veröffentlicht: (2025)
Advanced Long-Content Speech Recognition With Factorized Neural Transducer
von: Gong, Xun, et al.
Veröffentlicht: (2024)
von: Gong, Xun, et al.
Veröffentlicht: (2024)
Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams
von: He, Xiluo, et al.
Veröffentlicht: (2025)
von: He, Xiluo, et al.
Veröffentlicht: (2025)
StreamAAD: Decoding Spatial Auditory Attention with a Streaming Architecture
von: Qiu, Zelin, et al.
Veröffentlicht: (2024)
von: Qiu, Zelin, et al.
Veröffentlicht: (2024)
TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASR
von: Kumar, Shashi, et al.
Veröffentlicht: (2024)
von: Kumar, Shashi, et al.
Veröffentlicht: (2024)
Efficient Scaling for LLM-based ASR
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
Lightweight Target-Speaker-Based Overlap Transcription for Practical Streaming ASR
von: Pražák, Aleš, et al.
Veröffentlicht: (2025)
von: Pražák, Aleš, et al.
Veröffentlicht: (2025)
Whistle: Data-Efficient Multilingual and Crosslingual Speech Recognition via Weakly Phonetic Supervision
von: Yusuyin, Saierdaer, et al.
Veröffentlicht: (2024)
von: Yusuyin, Saierdaer, et al.
Veröffentlicht: (2024)
TDT-KWS: Fast And Accurate Keyword Spotting Using Token-and-duration Transducer
von: Xi, Yu, et al.
Veröffentlicht: (2024)
von: Xi, Yu, et al.
Veröffentlicht: (2024)
ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription
von: Le, Khanh, et al.
Veröffentlicht: (2025)
von: Le, Khanh, et al.
Veröffentlicht: (2025)
Pushing the Limits of Beam Search Decoding for Transducer-based ASR models
von: Grigoryan, Lilit, et al.
Veröffentlicht: (2025)
von: Grigoryan, Lilit, et al.
Veröffentlicht: (2025)
BrainWhisperer: Leveraging Large-Scale ASR Models for Neural Speech Decoding
von: Boccato, Tommaso, et al.
Veröffentlicht: (2026)
von: Boccato, Tommaso, et al.
Veröffentlicht: (2026)
Boosting Hybrid Autoregressive Transducer-based ASR with Internal Acoustic Model Training and Dual Blank Thresholding
von: Moriya, Takafumi, et al.
Veröffentlicht: (2024)
von: Moriya, Takafumi, et al.
Veröffentlicht: (2024)
Incremental FastPitch: Chunk-based High Quality Text to Speech
von: Du, Muyang, et al.
Veröffentlicht: (2024)
von: Du, Muyang, et al.
Veröffentlicht: (2024)
Unifying Streaming and Non-streaming Zipformer-based ASR
von: Sharma, Bidisha, et al.
Veröffentlicht: (2025)
von: Sharma, Bidisha, et al.
Veröffentlicht: (2025)
Decoding Strategies for Diffusion-Based ASR: A Systematic Evaluation of Confidence-Based Thresholding
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2026)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2026)
Emotion Neural Transducer for Fine-Grained Speech Emotion Recognition
von: Shen, Siyuan, et al.
Veröffentlicht: (2024)
von: Shen, Siyuan, et al.
Veröffentlicht: (2024)
Mamba for Streaming ASR Combined with Unimodal Aggregation
von: Fang, Ying, et al.
Veröffentlicht: (2024)
von: Fang, Ying, et al.
Veröffentlicht: (2024)
Phonemes vs. Projectors: An Investigation of Speech-Language Interfaces for LLM-based ASR
von: Li, Ziwei, et al.
Veröffentlicht: (2026)
von: Li, Ziwei, et al.
Veröffentlicht: (2026)
StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding
von: Guo, Dake, et al.
Veröffentlicht: (2025)
von: Guo, Dake, et al.
Veröffentlicht: (2025)
Speech Emotion Recognition with ASR Integration
von: Li, Yuanchao
Veröffentlicht: (2026)
von: Li, Yuanchao
Veröffentlicht: (2026)
Joint Beam Search Integrating CTC, Attention, and Transducer Decoders
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
Enhanced Hybrid Transducer and Attention Encoder Decoder with Text Data
von: Tang, Yun, et al.
Veröffentlicht: (2025)
von: Tang, Yun, et al.
Veröffentlicht: (2025)
Loss Masking Is Not Needed in Decoder-only Transformer for Discrete-token-based ASR
von: Chen, Qian, et al.
Veröffentlicht: (2023)
von: Chen, Qian, et al.
Veröffentlicht: (2023)
dLLM-ASR: A Faster Diffusion LLM-based Framework for Speech Recognition
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
The USTC-NERCSLIP Systems for The ICMC-ASR Challenge
von: Wu, Minghui, et al.
Veröffentlicht: (2024)
von: Wu, Minghui, et al.
Veröffentlicht: (2024)
Breaking Down Power Barriers in On-Device Streaming ASR: Insights and Solutions
von: Li, Yang, et al.
Veröffentlicht: (2024)
von: Li, Yang, et al.
Veröffentlicht: (2024)
Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding
von: Zhou, Haoran, et al.
Veröffentlicht: (2025)
von: Zhou, Haoran, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Masked Self-distilled Transducer-based Keyword Spotting with Semi-autoregressive Decoding
von: Xi, Yu, et al.
Veröffentlicht: (2025) -
SSCFormer: Push the Limit of Chunk-wise Conformer for Streaming ASR Using Sequentially Sampled Chunks and Chunked Causal Convolution
von: Wang, Fangyuan, et al.
Veröffentlicht: (2022) -
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
von: Thorbecke, Iuliia, et al.
Veröffentlicht: (2024) -
Chunked Attention-based Encoder-Decoder Model for Streaming Speech Recognition
von: Zeineldeen, Mohammad, et al.
Veröffentlicht: (2023) -
VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech
von: Du, Chenpeng, et al.
Veröffentlicht: (2024)