EffectiveASR: A Single-Step Non-Autoregressive Mandarin Speech Recognition Architecture with High Accuracy and Inference Speed
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhuang, Ziyang, Miao, Chenfeng, Zou, Kun, Fang, Ming, Wei, Tao, Li, Zijian, Cheng, Ning, Hu, Wei, Wang, Shaojun, Xiao, Jing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MDM-ASR: Bridging Accuracy and Efficiency in ASR with Diffusion-Based Non-Autoregressive Decoding
von: Yen, Hao, et al.
Veröffentlicht: (2026)
von: Yen, Hao, et al.
Veröffentlicht: (2026)
FireRedASR: Open-Source Industrial-Grade Mandarin Speech Recognition Models from Encoder-Decoder to LLM Integration
von: Xu, Kai-Tuo, et al.
Veröffentlicht: (2025)
von: Xu, Kai-Tuo, et al.
Veröffentlicht: (2025)
Speech Emotion Recognition with ASR Integration
von: Li, Yuanchao
Veröffentlicht: (2026)
von: Li, Yuanchao
Veröffentlicht: (2026)
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
Findings of the 2024 Mandarin Stuttering Event Detection and Automatic Speech Recognition Challenge
von: Xue, Hongfei, et al.
Veröffentlicht: (2024)
von: Xue, Hongfei, et al.
Veröffentlicht: (2024)
Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
von: Bai, Ye, et al.
Veröffentlicht: (2024)
von: Bai, Ye, et al.
Veröffentlicht: (2024)
Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding
von: Lin, Zijian, et al.
Veröffentlicht: (2025)
von: Lin, Zijian, et al.
Veröffentlicht: (2025)
LA-RAG:Enhancing LLM-based ASR Accuracy with Retrieval-Augmented Generation
von: Li, Shaojun, et al.
Veröffentlicht: (2024)
von: Li, Shaojun, et al.
Veröffentlicht: (2024)
SPG-Codec: Exploring the Role and Boundaries of Semantic Priors in Ultra-Low-Bitrate Neural Speech Coding
von: Zhao, Mingyu, et al.
Veröffentlicht: (2026)
von: Zhao, Mingyu, et al.
Veröffentlicht: (2026)
SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding
von: Wei, Linye, et al.
Veröffentlicht: (2025)
von: Wei, Linye, et al.
Veröffentlicht: (2025)
Towards a Single ASR Model That Generalizes to Disordered Speech
von: Tobin, Jimmy, et al.
Veröffentlicht: (2024)
von: Tobin, Jimmy, et al.
Veröffentlicht: (2024)
FairASR: Fair Audio Contrastive Learning for Automatic Speech Recognition
von: Kim, Jongsuk, et al.
Veröffentlicht: (2025)
von: Kim, Jongsuk, et al.
Veröffentlicht: (2025)
EfficientASR: Speech Recognition Network Compression via Attention Redundancy and Chunk-Level FFN Optimization
von: Wang, Jianzong, et al.
Veröffentlicht: (2024)
von: Wang, Jianzong, et al.
Veröffentlicht: (2024)
WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark
von: Ma, Linhan, et al.
Veröffentlicht: (2024)
von: Ma, Linhan, et al.
Veröffentlicht: (2024)
ChildMandarin: A Comprehensive Mandarin Speech Dataset for Young Children Aged 3-5
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
Investigation of Deep Neural Network Acoustic Modelling Approaches for Low Resource Accented Mandarin Speech Recognition
von: Xie, Xurong, et al.
Veröffentlicht: (2022)
von: Xie, Xurong, et al.
Veröffentlicht: (2022)
A Comprehensive Study on the Effectiveness of ASR Representations for Noise-Robust Speech Emotion Recognition
von: Shi, Xiaohan, et al.
Veröffentlicht: (2023)
von: Shi, Xiaohan, et al.
Veröffentlicht: (2023)
Elevating Robust Multi-Talker ASR by Decoupling Speaker Separation and Speech Recognition
von: Yang, Yufeng, et al.
Veröffentlicht: (2025)
von: Yang, Yufeng, et al.
Veröffentlicht: (2025)
Multi-Channel Differential ASR for Robust Wearer Speech Recognition on Smart Glasses
von: Yang, Yufeng, et al.
Veröffentlicht: (2025)
von: Yang, Yufeng, et al.
Veröffentlicht: (2025)
Efficient Scaling for LLM-based ASR
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
FireRedASR2S: A State-of-the-Art Industrial-Grade All-in-One Automatic Speech Recognition System
von: Xu, Kaituo, et al.
Veröffentlicht: (2026)
von: Xu, Kaituo, et al.
Veröffentlicht: (2026)
dLLM-ASR: A Faster Diffusion LLM-based Framework for Speech Recognition
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
ASR for Affective Speech: Investigating Impact of Emotion and Speech Generative Strategy
von: Wu, Ya-Tse, et al.
Veröffentlicht: (2026)
von: Wu, Ya-Tse, et al.
Veröffentlicht: (2026)
AS-ASR: A Lightweight Framework for Aphasia-Specific Automatic Speech Recognition
von: Bao, Chen, et al.
Veröffentlicht: (2025)
von: Bao, Chen, et al.
Veröffentlicht: (2025)
Investigating Zero-Shot Generalizability on Mandarin-English Code-Switched ASR and Speech-to-text Translation of Recent Foundation Models with Self-Supervision and Weak Supervision
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2023)
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2023)
Conversational Speech Recognition by Learning Audio-textual Cross-modal Contextual Representation
von: Wei, Kun, et al.
Veröffentlicht: (2023)
von: Wei, Kun, et al.
Veröffentlicht: (2023)
MedASR: An Open-Source Model for High-Accuracy Medical Dictation
von: Wu, Ke, et al.
Veröffentlicht: (2026)
von: Wu, Ke, et al.
Veröffentlicht: (2026)
AdaLTM: Adaptive Layer-wise Task Vector Merging for Categorical Speech Emotion Recognition with ASR Knowledge Integration
von: Lee, Chia-Yu, et al.
Veröffentlicht: (2026)
von: Lee, Chia-Yu, et al.
Veröffentlicht: (2026)
Selective Invocation for Multilingual ASR: A Cost-effective Approach Adapting to Speech Recognition Difficulty
von: Xue, Hongfei, et al.
Veröffentlicht: (2025)
von: Xue, Hongfei, et al.
Veröffentlicht: (2025)
Towards Robust Dysarthric Speech Recognition: LLM-Agent Post-ASR Correction Beyond WER
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2026)
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2026)
Speaker-Smoothed kNN Speaker Adaptation for End-to-End ASR
von: Li, Shaojun, et al.
Veröffentlicht: (2024)
von: Li, Shaojun, et al.
Veröffentlicht: (2024)
MNV-17: A High-Quality Performative Mandarin Dataset for Nonverbal Vocalization Recognition in Speech
von: Mai, Jialong, et al.
Veröffentlicht: (2025)
von: Mai, Jialong, et al.
Veröffentlicht: (2025)
Mixture of LoRA Experts with Multi-Modal and Multi-Granularity LLM Generative Error Correction for Accented Speech Recognition
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
DSFlow: Dual Supervision and Step-Aware Architecture for One-Step Flow Matching Speech Synthesis
von: Lin, Bin, et al.
Veröffentlicht: (2026)
von: Lin, Bin, et al.
Veröffentlicht: (2026)
ICMC-ASR: The ICASSP 2024 In-Car Multi-Channel Automatic Speech Recognition Challenge
von: Wang, He, et al.
Veröffentlicht: (2024)
von: Wang, He, et al.
Veröffentlicht: (2024)
VINP: Variational Bayesian Inference with Neural Speech Prior for Joint ASR-Effective Speech Dereverberation and Blind RIR Identification
von: Wang, Pengyu, et al.
Veröffentlicht: (2025)
von: Wang, Pengyu, et al.
Veröffentlicht: (2025)
Temporal-Frequency State Space Duality: An Efficient Paradigm for Speech Emotion Recognition
von: Zhao, Jiaqi, et al.
Veröffentlicht: (2024)
von: Zhao, Jiaqi, et al.
Veröffentlicht: (2024)
Unifying Speech Recognition, Synthesis and Conversion with Autoregressive Transformers
von: Cai, Runyuan, et al.
Veröffentlicht: (2026)
von: Cai, Runyuan, et al.
Veröffentlicht: (2026)
SW-ASR: A Context-Aware Hybrid ASR Pipeline for Robust Single Word Speech Recognition
von: Sharma, Manali, et al.
Veröffentlicht: (2026)
von: Sharma, Manali, et al.
Veröffentlicht: (2026)
MaLa-ASR: Multimedia-Assisted LLM-Based ASR
von: Yang, Guanrou, et al.
Veröffentlicht: (2024)
von: Yang, Guanrou, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MDM-ASR: Bridging Accuracy and Efficiency in ASR with Diffusion-Based Non-Autoregressive Decoding
von: Yen, Hao, et al.
Veröffentlicht: (2026) -
FireRedASR: Open-Source Industrial-Grade Mandarin Speech Recognition Models from Encoder-Decoder to LLM Integration
von: Xu, Kai-Tuo, et al.
Veröffentlicht: (2025) -
Speech Emotion Recognition with ASR Integration
von: Li, Yuanchao
Veröffentlicht: (2026) -
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
von: Wang, Yujin, et al.
Veröffentlicht: (2022) -
Findings of the 2024 Mandarin Stuttering Event Detection and Automatic Speech Recognition Challenge
von: Xue, Hongfei, et al.
Veröffentlicht: (2024)