Paraformer-v2: An improved non-autoregressive transformer for noise-robust speech recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | An, Keyu, Li, Zerui, Gao, Zhifu, Zhang, Shiliang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards noise-robust speech inversion through multi-task learning with speech enhancement
von: Tabatabaee, Saba, et al.
Veröffentlicht: (2026)
von: Tabatabaee, Saba, et al.
Veröffentlicht: (2026)
From the perspective of perceptual speech quality: The robustness of frequency bands to noise
von: Fan, Junyi, et al.
Veröffentlicht: (2025)
von: Fan, Junyi, et al.
Veröffentlicht: (2025)
CR-CTC: Consistency regularization on CTC for improved speech recognition
von: Yao, Zengwei, et al.
Veröffentlicht: (2024)
von: Yao, Zengwei, et al.
Veröffentlicht: (2024)
Are Transformers in Pre-trained LM A Good ASR Encoder? An Empirical Study
von: An, Keyu, et al.
Veröffentlicht: (2024)
von: An, Keyu, et al.
Veröffentlicht: (2024)
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
von: Ducorroy, Alexandre, et al.
Veröffentlicht: (2025)
von: Ducorroy, Alexandre, et al.
Veröffentlicht: (2025)
Phoneme-based speech recognition driven by large language models and sampling marginalization
von: Ma, Te, et al.
Veröffentlicht: (2025)
von: Ma, Te, et al.
Veröffentlicht: (2025)
Graph-based multi-Feature fusion method for speech emotion recognition
von: Liu, Xueyu, et al.
Veröffentlicht: (2024)
von: Liu, Xueyu, et al.
Veröffentlicht: (2024)
Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition
von: Zhang, Yiru, et al.
Veröffentlicht: (2025)
von: Zhang, Yiru, et al.
Veröffentlicht: (2025)
Improving child speech recognition with augmented child-like speech
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024)
Language model integration based on memory control for sequence to sequence speech recognition
von: Cho, Jaejin, et al.
Veröffentlicht: (2018)
von: Cho, Jaejin, et al.
Veröffentlicht: (2018)
A lightweight and robust method for blind wideband-to-fullband extension of speech
von: Büthe, Jan, et al.
Veröffentlicht: (2024)
von: Büthe, Jan, et al.
Veröffentlicht: (2024)
CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
von: Du, Zhihao, et al.
Veröffentlicht: (2024)
von: Du, Zhihao, et al.
Veröffentlicht: (2024)
Explore the Reinforcement Learning for the LLM based ASR and TTS system
von: Gao, Changfeng, et al.
Veröffentlicht: (2025)
von: Gao, Changfeng, et al.
Veröffentlicht: (2025)
GDiffuSE: Diffusion-based speech enhancement with noise model guidance
von: Yanir, Efrayim, et al.
Veröffentlicht: (2025)
von: Yanir, Efrayim, et al.
Veröffentlicht: (2025)
Charting 15 years of progress in deep learning for speech emotion recognition: A replication study
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2025)
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2025)
PROCTER: PROnunciation-aware ConTextual adaptER for personalized speech recognition in neural transducers
von: Pandey, Rahul, et al.
Veröffentlicht: (2023)
von: Pandey, Rahul, et al.
Veröffentlicht: (2023)
Adaptive ship-radiated noise recognition with learnable fine-grained wavelet transform
von: Xie, Yuan, et al.
Veröffentlicht: (2023)
von: Xie, Yuan, et al.
Veröffentlicht: (2023)
Explainable speech emotion recognition through attentive pooling: insights from attention-based temporal localization
von: Leygue, Tahitoa, et al.
Veröffentlicht: (2025)
von: Leygue, Tahitoa, et al.
Veröffentlicht: (2025)
AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection
von: Gong, Rong, et al.
Veröffentlicht: (2024)
von: Gong, Rong, et al.
Veröffentlicht: (2024)
LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
von: Dong, Lukuang, et al.
Veröffentlicht: (2026)
von: Dong, Lukuang, et al.
Veröffentlicht: (2026)
Effects of automotive microphone frequency response characteristics and noise conditions on speech and ASR quality -- an experimental evaluation
von: Buccoli, Michele, et al.
Veröffentlicht: (2025)
von: Buccoli, Michele, et al.
Veröffentlicht: (2025)
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception
von: Zhang, Jiawei, et al.
Veröffentlicht: (2024)
von: Zhang, Jiawei, et al.
Veröffentlicht: (2024)
Heterogeneous bimodal attention fusion for speech emotion recognition
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
EBEN: Extreme bandwidth extension network applied to speech signals captured with noise-resilient body-conduction microphones
von: Hauret, Julien, et al.
Veröffentlicht: (2022)
von: Hauret, Julien, et al.
Veröffentlicht: (2022)
Toward noise-robust whisper keyword spotting on headphones with in-earcup microphone and curriculum learning
von: Yang, Qiaoyu
Veröffentlicht: (2025)
von: Yang, Qiaoyu
Veröffentlicht: (2025)
Tweaking autoregressive methods for inpainting of gaps in audio signals
von: Mokrý, Ondřej, et al.
Veröffentlicht: (2024)
von: Mokrý, Ondřej, et al.
Veröffentlicht: (2024)
Advancing automatic speech recognition using feature fusion with self-supervised learning features: A case study on Fearless Steps Apollo corpus
von: Chen, Szu-Jui, et al.
Veröffentlicht: (2026)
von: Chen, Szu-Jui, et al.
Veröffentlicht: (2026)
Towards robust paralinguistic assessment for real-world mobile health (mHealth) monitoring: an initial study of reverberation effects on speech
von: Dineley, Judith, et al.
Veröffentlicht: (2023)
von: Dineley, Judith, et al.
Veröffentlicht: (2023)
Masked Self-distilled Transducer-based Keyword Spotting with Semi-autoregressive Decoding
von: Xi, Yu, et al.
Veröffentlicht: (2025)
von: Xi, Yu, et al.
Veröffentlicht: (2025)
Regularized autoregressive modeling and its application to audio signal reconstruction
von: Mokrý, Ondřej, et al.
Veröffentlicht: (2024)
von: Mokrý, Ondřej, et al.
Veröffentlicht: (2024)
Self-consistent context aware conformer transducer for speech recognition
von: Kolokolov, Konstantin, et al.
Veröffentlicht: (2024)
von: Kolokolov, Konstantin, et al.
Veröffentlicht: (2024)
Enhancing CTC-based speech recognition with diverse modeling units
von: Han, Shiyi, et al.
Veröffentlicht: (2024)
von: Han, Shiyi, et al.
Veröffentlicht: (2024)
An efficient text augmentation approach for contextualized Mandarin speech recognition
von: Zheng, Naijun, et al.
Veröffentlicht: (2024)
von: Zheng, Naijun, et al.
Veröffentlicht: (2024)
Zipformer: A faster and better encoder for automatic speech recognition
von: Yao, Zengwei, et al.
Veröffentlicht: (2023)
von: Yao, Zengwei, et al.
Veröffentlicht: (2023)
LLM-based phoneme-to-grapheme for phoneme-based speech recognition
von: Ma, Te, et al.
Veröffentlicht: (2025)
von: Ma, Te, et al.
Veröffentlicht: (2025)
Robustifying automatic speech recognition by extracting slowly varying features
von: Pizarro, Matías, et al.
Veröffentlicht: (2021)
von: Pizarro, Matías, et al.
Veröffentlicht: (2021)
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
von: Maiti, Soumi, et al.
Veröffentlicht: (2023)
von: Maiti, Soumi, et al.
Veröffentlicht: (2023)
On the effectiveness of enrollment speech augmentation for Target Speaker Extraction
von: Li, Junjie, et al.
Veröffentlicht: (2024)
von: Li, Junjie, et al.
Veröffentlicht: (2024)
Introduction to speech recognition
von: Dauphin, Gabriel
Veröffentlicht: (2024)
von: Dauphin, Gabriel
Veröffentlicht: (2024)
Ähnliche Einträge
-
Towards noise-robust speech inversion through multi-task learning with speech enhancement
von: Tabatabaee, Saba, et al.
Veröffentlicht: (2026) -
From the perspective of perceptual speech quality: The robustness of frequency bands to noise
von: Fan, Junyi, et al.
Veröffentlicht: (2025) -
CR-CTC: Consistency regularization on CTC for improved speech recognition
von: Yao, Zengwei, et al.
Veröffentlicht: (2024) -
Are Transformers in Pre-trained LM A Good ASR Encoder? An Empirical Study
von: An, Keyu, et al.
Veröffentlicht: (2024) -
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
von: Ducorroy, Alexandre, et al.
Veröffentlicht: (2025)