Qifusion-Net: Layer-adapted Stream/Non-stream Model for End-to-End Multi-Accent Speech Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Jinming, Fang, Jingyi, Zheng, Yuanzhong, Wang, Yaoxuan, Fei, Haojun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
QvTAD: Differential Relative Attribute Learning for Voice Timbre Attribute Detection
von: Wu, Zhiyu, et al.
Veröffentlicht: (2025)
von: Wu, Zhiyu, et al.
Veröffentlicht: (2025)
Qieemo: Speech Is All You Need in the Emotion Recognition in Conversations
von: Chen, Jinming, et al.
Veröffentlicht: (2025)
von: Chen, Jinming, et al.
Veröffentlicht: (2025)
SpecWav-Attack: Leveraging Spectrogram Resizing and Wav2Vec 2.0 for Attacking Anonymized Speech
von: Li, Yuqi, et al.
Veröffentlicht: (2025)
von: Li, Yuqi, et al.
Veröffentlicht: (2025)
Multi-Scale Accent Modeling and Disentangling for Multi-Speaker Multi-Accent Text-to-Speech Synthesis
von: Zhou, Xuehao, et al.
Veröffentlicht: (2024)
von: Zhou, Xuehao, et al.
Veröffentlicht: (2024)
MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition
von: Mu, Bingshen, et al.
Veröffentlicht: (2024)
von: Mu, Bingshen, et al.
Veröffentlicht: (2024)
Interpreting End-to-End Deep Learning Models for Speech Source Localization Using Layer-wise Relevance Propagation
von: Comanducci, Luca, et al.
Veröffentlicht: (2024)
von: Comanducci, Luca, et al.
Veröffentlicht: (2024)
Speaker Adaptation for Quantised End-to-End ASR Models
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
IKFST: IOO and KOO Algorithms for Accelerated and Precise WFST-based End-to-End Automatic Speech Recognition
von: Zhuang, Zhuoran, et al.
Veröffentlicht: (2026)
von: Zhuang, Zhuoran, et al.
Veröffentlicht: (2026)
MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy Speech
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
End-to-End Target Speaker Speech Recognition Using Context-Aware Attention Mechanisms for Challenging Enrollment Scenario
von: Ghane, Mohsen, et al.
Veröffentlicht: (2025)
von: Ghane, Mohsen, et al.
Veröffentlicht: (2025)
End-to-End Integration of Speech Emotion Recognition with Voice Activity Detection using Self-Supervised Learning Features
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2024)
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2024)
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
von: Nespoli, Francesco, et al.
Veröffentlicht: (2024)
von: Nespoli, Francesco, et al.
Veröffentlicht: (2024)
An Efficient End-to-End Approach to Noise Invariant Speech Features via Multi-Task Learning
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2024)
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2024)
Post-decoder Biasing for End-to-End Speech Recognition of Multi-turn Medical Interview
von: Liu, Heyang, et al.
Veröffentlicht: (2024)
von: Liu, Heyang, et al.
Veröffentlicht: (2024)
CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
von: Chen, Junyang, et al.
Veröffentlicht: (2026)
von: Chen, Junyang, et al.
Veröffentlicht: (2026)
Mixture of LoRA Experts with Multi-Modal and Multi-Granularity LLM Generative Error Correction for Accented Speech Recognition
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
GE2E-AC: Generalized End-to-End Loss Training for Accent Classification
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2024)
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2024)
Improving Practical Aspects of End-to-End Multi-Talker Speech Recognition for Online and Offline Scenarios
von: Subramanian, Aswin Shanmugam, et al.
Veröffentlicht: (2025)
von: Subramanian, Aswin Shanmugam, et al.
Veröffentlicht: (2025)
Investigation of Deep Neural Network Acoustic Modelling Approaches for Low Resource Accented Mandarin Speech Recognition
von: Xie, Xurong, et al.
Veröffentlicht: (2022)
von: Xie, Xurong, et al.
Veröffentlicht: (2022)
SAML: Speaker Adaptive Mixture of LoRA Experts for End-to-End ASR
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
WMCodec: End-to-End Neural Speech Codec with Deep Watermarking for Authenticity Verification
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
Central Kurdish Text-to-Speech Synthesis with Novel End-to-End Transformer Training
von: Ahmad, Hawraz A., et al.
Veröffentlicht: (2024)
von: Ahmad, Hawraz A., et al.
Veröffentlicht: (2024)
Time and Tokens: Benchmarking End-to-End Speech Dysfluency Detection
von: Zhou, Xuanru, et al.
Veröffentlicht: (2024)
von: Zhou, Xuanru, et al.
Veröffentlicht: (2024)
SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion
von: Guo, Zhao, et al.
Veröffentlicht: (2025)
von: Guo, Zhao, et al.
Veröffentlicht: (2025)
LS-EEND: Long-Form Streaming End-to-End Neural Diarization with Online Attractor Extraction
von: Liang, Di, et al.
Veröffentlicht: (2024)
von: Liang, Di, et al.
Veröffentlicht: (2024)
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
von: Guo, Yinlin, et al.
Veröffentlicht: (2024)
von: Guo, Yinlin, et al.
Veröffentlicht: (2024)
Magnetoencephalography (MEG) Based Non-Invasive Chinese Speech Decoding
von: Jia, Zhihong, et al.
Veröffentlicht: (2025)
von: Jia, Zhihong, et al.
Veröffentlicht: (2025)
Clustering and Mining Accented Speech for Inclusive and Fair Speech Recognition
von: Kim, Jaeyoung, et al.
Veröffentlicht: (2024)
von: Kim, Jaeyoung, et al.
Veröffentlicht: (2024)
Neural Scoring: A Refreshed End-to-End Approach for Speaker Recognition in Complex Conditions
von: Lin, Wan, et al.
Veröffentlicht: (2024)
von: Lin, Wan, et al.
Veröffentlicht: (2024)
Streaming Non-Autoregressive Model for Accent Conversion and Pronunciation Improvement
von: Nguyen, Tuan-Nam, et al.
Veröffentlicht: (2025)
von: Nguyen, Tuan-Nam, et al.
Veröffentlicht: (2025)
Speech-to-See: End-to-End Speech-Driven Open-Set Object Detection
von: Lu, Wenhuan, et al.
Veröffentlicht: (2025)
von: Lu, Wenhuan, et al.
Veröffentlicht: (2025)
StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
Harnessing the Zero-Shot Power of Instruction-Tuned Large Language Model in End-to-End Speech Recognition
von: Higuchi, Yosuke, et al.
Veröffentlicht: (2023)
von: Higuchi, Yosuke, et al.
Veröffentlicht: (2023)
Code-Switching in End-to-End Automatic Speech Recognition: A Systematic Literature Review
von: Agro, Maha Tufail, et al.
Veröffentlicht: (2025)
von: Agro, Maha Tufail, et al.
Veröffentlicht: (2025)
On Improving Error Resilience of Neural End-to-End Speech Coders
von: Gupta, Kishan, et al.
Veröffentlicht: (2024)
von: Gupta, Kishan, et al.
Veröffentlicht: (2024)
Gammatonegram Representation for End-to-End Dysarthric Speech Processing Tasks: Speech Recognition, Speaker Identification, and Intelligibility Assessment
von: Farhadipour, Aref, et al.
Veröffentlicht: (2023)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2023)
Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
von: Li, Tianpeng, et al.
Veröffentlicht: (2025)
von: Li, Tianpeng, et al.
Veröffentlicht: (2025)
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition
von: Hirano, Yuta, et al.
Veröffentlicht: (2025)
von: Hirano, Yuta, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
QvTAD: Differential Relative Attribute Learning for Voice Timbre Attribute Detection
von: Wu, Zhiyu, et al.
Veröffentlicht: (2025) -
Qieemo: Speech Is All You Need in the Emotion Recognition in Conversations
von: Chen, Jinming, et al.
Veröffentlicht: (2025) -
SpecWav-Attack: Leveraging Spectrogram Resizing and Wav2Vec 2.0 for Attacking Anonymized Speech
von: Li, Yuqi, et al.
Veröffentlicht: (2025) -
Multi-Scale Accent Modeling and Disentangling for Multi-Speaker Multi-Accent Text-to-Speech Synthesis
von: Zhou, Xuehao, et al.
Veröffentlicht: (2024) -
MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition
von: Mu, Bingshen, et al.
Veröffentlicht: (2024)