Zipformer: A faster and better encoder for automatic speech recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yao, Zengwei, Guo, Liyong, Yang, Xiaoyu, Kang, Wei, Kuang, Fangjun, Yang, Yifan, Jin, Zengrui, Lin, Long, Povey, Daniel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CR-CTC: Consistency regularization on CTC for improved speech recognition
von: Yao, Zengwei, et al.
Veröffentlicht: (2024)
von: Yao, Zengwei, et al.
Veröffentlicht: (2024)
Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context
von: Kang, Wei, et al.
Veröffentlicht: (2023)
von: Kang, Wei, et al.
Veröffentlicht: (2023)
PromptASR for contextualized ASR with controllable style
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2023)
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2023)
ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
von: Zhu, Han, et al.
Veröffentlicht: (2025)
von: Zhu, Han, et al.
Veröffentlicht: (2025)
LibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization
von: Jin, Zengrui, et al.
Veröffentlicht: (2024)
von: Jin, Zengrui, et al.
Veröffentlicht: (2024)
k2SSL: A Faster and Better Framework for Self-Supervised Speech Representation Learning
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition
von: Zhang, Yiru, et al.
Veröffentlicht: (2025)
von: Zhang, Yiru, et al.
Veröffentlicht: (2025)
Robustifying automatic speech recognition by extracting slowly varying features
von: Pizarro, Matías, et al.
Veröffentlicht: (2021)
von: Pizarro, Matías, et al.
Veröffentlicht: (2021)
AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection
von: Gong, Rong, et al.
Veröffentlicht: (2024)
von: Gong, Rong, et al.
Veröffentlicht: (2024)
Training dynamic models using early exits for automatic speech recognition on resource-constrained devices
von: Wright, George August, et al.
Veröffentlicht: (2023)
von: Wright, George August, et al.
Veröffentlicht: (2023)
Distilling a speech and music encoder with task arithmetic
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2025)
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2025)
Flow2GAN: Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-step High-Fidelity Audio Generation
von: Yao, Zengwei, et al.
Veröffentlicht: (2025)
von: Yao, Zengwei, et al.
Veröffentlicht: (2025)
Graph-based multi-Feature fusion method for speech emotion recognition
von: Liu, Xueyu, et al.
Veröffentlicht: (2024)
von: Liu, Xueyu, et al.
Veröffentlicht: (2024)
OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models
von: Zhu, Han, et al.
Veröffentlicht: (2026)
von: Zhu, Han, et al.
Veröffentlicht: (2026)
SeMaScore : a new evaluation metric for automatic speech recognition tasks
von: Sasindran, Zitha, et al.
Veröffentlicht: (2024)
von: Sasindran, Zitha, et al.
Veröffentlicht: (2024)
Automated speech audiometry: Can it work using open-source pre-trained Kaldi-NL automatic speech recognition?
von: Araiza-Illan, Gloria, et al.
Veröffentlicht: (2023)
von: Araiza-Illan, Gloria, et al.
Veröffentlicht: (2023)
Advancing automatic speech recognition using feature fusion with self-supervised learning features: A case study on Fearless Steps Apollo corpus
von: Chen, Szu-Jui, et al.
Veröffentlicht: (2026)
von: Chen, Szu-Jui, et al.
Veröffentlicht: (2026)
ZipEnhancer: Dual-Path Down-Up Sampling-based Zipformer for Monaural Speech Enhancement
von: Wang, Haoxu, et al.
Veröffentlicht: (2025)
von: Wang, Haoxu, et al.
Veröffentlicht: (2025)
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
von: Ducorroy, Alexandre, et al.
Veröffentlicht: (2025)
von: Ducorroy, Alexandre, et al.
Veröffentlicht: (2025)
Unifying Streaming and Non-streaming Zipformer-based ASR
von: Sharma, Bidisha, et al.
Veröffentlicht: (2025)
von: Sharma, Bidisha, et al.
Veröffentlicht: (2025)
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
von: Maiti, Soumi, et al.
Veröffentlicht: (2023)
von: Maiti, Soumi, et al.
Veröffentlicht: (2023)
Language model integration based on memory control for sequence to sequence speech recognition
von: Cho, Jaejin, et al.
Veröffentlicht: (2018)
von: Cho, Jaejin, et al.
Veröffentlicht: (2018)
Phoneme-based speech recognition driven by large language models and sampling marginalization
von: Ma, Te, et al.
Veröffentlicht: (2025)
von: Ma, Te, et al.
Veröffentlicht: (2025)
FreeCodec: A disentangled neural speech codec with fewer tokens
von: Zheng, Youqiang, et al.
Veröffentlicht: (2024)
von: Zheng, Youqiang, et al.
Veröffentlicht: (2024)
Heterogeneous bimodal attention fusion for speech emotion recognition
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
Improving child speech recognition with augmented child-like speech
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024)
An automatic mixing speech enhancement system for multi-track audio
von: Liu, Xiaojing, et al.
Veröffentlicht: (2024)
von: Liu, Xiaojing, et al.
Veröffentlicht: (2024)
Paraformer-v2: An improved non-autoregressive transformer for noise-robust speech recognition
von: An, Keyu, et al.
Veröffentlicht: (2024)
von: An, Keyu, et al.
Veröffentlicht: (2024)
Charting 15 years of progress in deep learning for speech emotion recognition: A replication study
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2025)
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2025)
PROCTER: PROnunciation-aware ConTextual adaptER for personalized speech recognition in neural transducers
von: Pandey, Rahul, et al.
Veröffentlicht: (2023)
von: Pandey, Rahul, et al.
Veröffentlicht: (2023)
Explainable speech emotion recognition through attentive pooling: insights from attention-based temporal localization
von: Leygue, Tahitoa, et al.
Veröffentlicht: (2025)
von: Leygue, Tahitoa, et al.
Veröffentlicht: (2025)
A correlation-permutation approach for speech-music encoders model merging
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2025)
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2025)
DualSep: A Light-weight dual-encoder convolutional recurrent network for real-time in-car speech separation
von: Wang, Ziqian, et al.
Veröffentlicht: (2024)
von: Wang, Ziqian, et al.
Veröffentlicht: (2024)
Revisiting speech segmentation and lexicon learning with better features
von: Kamper, Herman, et al.
Veröffentlicht: (2024)
von: Kamper, Herman, et al.
Veröffentlicht: (2024)
LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
Introduction to speech recognition
von: Dauphin, Gabriel
Veröffentlicht: (2024)
von: Dauphin, Gabriel
Veröffentlicht: (2024)
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
von: Dong, Lukuang, et al.
Veröffentlicht: (2026)
von: Dong, Lukuang, et al.
Veröffentlicht: (2026)
Self-consistent context aware conformer transducer for speech recognition
von: Kolokolov, Konstantin, et al.
Veröffentlicht: (2024)
von: Kolokolov, Konstantin, et al.
Veröffentlicht: (2024)
LLM-based phoneme-to-grapheme for phoneme-based speech recognition
von: Ma, Te, et al.
Veröffentlicht: (2025)
von: Ma, Te, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CR-CTC: Consistency regularization on CTC for improved speech recognition
von: Yao, Zengwei, et al.
Veröffentlicht: (2024) -
Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context
von: Kang, Wei, et al.
Veröffentlicht: (2023) -
PromptASR for contextualized ASR with controllable style
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2023) -
ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
von: Zhu, Han, et al.
Veröffentlicht: (2025) -
LibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization
von: Jin, Zengrui, et al.
Veröffentlicht: (2024)