An Effective Mixture-Of-Experts Approach For Code-Switching Speech Recognition Leveraging Encoder Disentanglement
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Tzu-Ting, Wang, Hsin-Wei, Wang, Yi-Cheng, Lin, Chi-Han, Chen, Berlin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enhancing Code-Switching ASR Leveraging Non-Peaky CTC Loss and Deep Language Posterior Injection
von: Yang, Tzu-Ting, et al.
Veröffentlicht: (2024)
von: Yang, Tzu-Ting, et al.
Veröffentlicht: (2024)
An Effective Context-Balanced Adaptation Approach for Long-Tailed Speech Recognition
von: Wang, Yi-Cheng, et al.
Veröffentlicht: (2024)
von: Wang, Yi-Cheng, et al.
Veröffentlicht: (2024)
CAMEL: Cross-Attention Enhanced Mixture-of-Experts and Language Bias for Code-Switching Speech Recognition
von: Wang, He, et al.
Veröffentlicht: (2024)
von: Wang, He, et al.
Veröffentlicht: (2024)
Enhancing Code-Switching Speech Recognition with LID-Based Collaborative Mixture of Experts Model
von: Huang, Hukai, et al.
Veröffentlicht: (2024)
von: Huang, Hukai, et al.
Veröffentlicht: (2024)
Boosting Code-Switching ASR with Mixture of Experts Enhanced Speech-Conditioned LLM
von: Zhang, Fengrun, et al.
Veröffentlicht: (2024)
von: Zhang, Fengrun, et al.
Veröffentlicht: (2024)
ConPCO: Preserving Phoneme Characteristics for Automatic Pronunciation Assessment Leveraging Contrastive Ordinal Regularization
von: Yan, Bi-Cheng, et al.
Veröffentlicht: (2024)
von: Yan, Bi-Cheng, et al.
Veröffentlicht: (2024)
Speech-Aware Neural Diarization with Encoder-Decoder Attractor Guided by Attention Constraints
von: Lee, PeiYing, et al.
Veröffentlicht: (2024)
von: Lee, PeiYing, et al.
Veröffentlicht: (2024)
Attention-Guided Adaptation for Code-Switching Speech Recognition
von: Aditya, Bobbi, et al.
Veröffentlicht: (2023)
von: Aditya, Bobbi, et al.
Veröffentlicht: (2023)
Effective Noise-aware Data Simulation for Domain-adaptive Speech Enhancement Leveraging Dynamic Stochastic Perturbation
von: Wang, Chien-Chun, et al.
Veröffentlicht: (2024)
von: Wang, Chien-Chun, et al.
Veröffentlicht: (2024)
Leveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech Enhancement
von: Ren, Wenze, et al.
Veröffentlicht: (2024)
von: Ren, Wenze, et al.
Veröffentlicht: (2024)
Universal Robust Speech Adaptation for Cross-Domain Speech Recognition and Enhancement
von: Wang, Chien-Chun, et al.
Veröffentlicht: (2026)
von: Wang, Chien-Chun, et al.
Veröffentlicht: (2026)
Disentangling Dual-Encoder Masked Autoencoder for Respiratory Sound Classification
von: Wei, Peidong, et al.
Veröffentlicht: (2025)
von: Wei, Peidong, et al.
Veröffentlicht: (2025)
Leveraging Mixture of Experts for Improved Speech Deepfake Detection
von: Negroni, Viola, et al.
Veröffentlicht: (2024)
von: Negroni, Viola, et al.
Veröffentlicht: (2024)
Variational Auto-Encoder Based Variability Encoding for Dysarthric Speech Recognition
von: Xie, Xurong, et al.
Veröffentlicht: (2022)
von: Xie, Xurong, et al.
Veröffentlicht: (2022)
An Effective Automated Speaking Assessment Approach to Mitigating Data Scarcity and Imbalanced Distribution
von: Lo, Tien-Hong, et al.
Veröffentlicht: (2024)
von: Lo, Tien-Hong, et al.
Veröffentlicht: (2024)
Mixture of LoRA Experts with Multi-Modal and Multi-Granularity LLM Generative Error Correction for Accented Speech Recognition
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy Speech
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
Noisy Disentanglement with Tri-stage Training for Noise-Robust Speech Recognition
von: Chen, Shuangyuan, et al.
Veröffentlicht: (2025)
von: Chen, Shuangyuan, et al.
Veröffentlicht: (2025)
Channel-Aware Domain-Adaptive Generative Adversarial Network for Robust Speech Recognition
von: Wang, Chien-Chun, et al.
Veröffentlicht: (2024)
von: Wang, Chien-Chun, et al.
Veröffentlicht: (2024)
Enhancing Speech Large Language Models with Prompt-Aware Mixture of Audio Encoders
von: Shan, Weiqiao, et al.
Veröffentlicht: (2025)
von: Shan, Weiqiao, et al.
Veröffentlicht: (2025)
Learning Expressive Disentangled Speech Representations with Soft Speech Units and Adversarial Style Augmentation
von: Deng, Yimin, et al.
Veröffentlicht: (2024)
von: Deng, Yimin, et al.
Veröffentlicht: (2024)
Rethinking Processing Distortions: Disentangling the Impact of Speech Enhancement Errors on Speech Recognition Performance
von: Ochiai, Tsubasa, et al.
Veröffentlicht: (2024)
von: Ochiai, Tsubasa, et al.
Veröffentlicht: (2024)
EMO-SUPERB: An In-depth Look at Speech Emotion Recognition
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation
von: Chen, Szu-Chi, et al.
Veröffentlicht: (2026)
von: Chen, Szu-Chi, et al.
Veröffentlicht: (2026)
TG-ASR: Translation-Guided Learning with Parallel Gated Cross Attention for Low-Resource Automatic Speech Recognition
von: Yang, Cheng-Yeh, et al.
Veröffentlicht: (2026)
von: Yang, Cheng-Yeh, et al.
Veröffentlicht: (2026)
EffectiveASR: A Single-Step Non-Autoregressive Mandarin Speech Recognition Architecture with High Accuracy and Inference Speed
von: Zhuang, Ziyang, et al.
Veröffentlicht: (2024)
von: Zhuang, Ziyang, et al.
Veröffentlicht: (2024)
Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
von: Huo, Mingyue, et al.
Veröffentlicht: (2025)
von: Huo, Mingyue, et al.
Veröffentlicht: (2025)
CS-Dialogue: A 104-Hour Dataset of Spontaneous Mandarin-English Code-Switching Dialogues for Speech Recognition
von: Zhou, Jiaming, et al.
Veröffentlicht: (2025)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2025)
SC-MoE: Switch Conformer Mixture of Experts for Unified Streaming and Non-streaming Code-Switching ASR
von: Ye, Shuaishuai, et al.
Veröffentlicht: (2024)
von: Ye, Shuaishuai, et al.
Veröffentlicht: (2024)
Self-Supervised Speech Quality Assessment (S3QA): Leveraging Speech Foundation Models for a Scalable Speech Quality Metric
von: Ogg, Mattson, et al.
Veröffentlicht: (2025)
von: Ogg, Mattson, et al.
Veröffentlicht: (2025)
Fx-Encoder++: Extracting Instrument-Wise Audio Effects Representations from Mixtures
von: Yeh, Yen-Tung, et al.
Veröffentlicht: (2025)
von: Yeh, Yen-Tung, et al.
Veröffentlicht: (2025)
Learning Disentangled Speech Representations with Contrastive Learning and Time-Invariant Retrieval
von: Deng, Yimin, et al.
Veröffentlicht: (2024)
von: Deng, Yimin, et al.
Veröffentlicht: (2024)
Leveraging Self-Supervised Models for Automatic Whispered Speech Recognition
von: Farhadipour, Aref, et al.
Veröffentlicht: (2024)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2024)
Leveraging Broadcast Media Subtitle Transcripts for Automatic Speech Recognition and Subtitling
von: Poncelet, Jakob, et al.
Veröffentlicht: (2025)
von: Poncelet, Jakob, et al.
Veröffentlicht: (2025)
Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
von: Liao, Shijia, et al.
Veröffentlicht: (2024)
von: Liao, Shijia, et al.
Veröffentlicht: (2024)
Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
von: Bai, Ye, et al.
Veröffentlicht: (2024)
von: Bai, Ye, et al.
Veröffentlicht: (2024)
A Self-Training Approach for Whisper to Enhance Long Dysarthric Speech Recognition
von: Wang, Shiyao, et al.
Veröffentlicht: (2025)
von: Wang, Shiyao, et al.
Veröffentlicht: (2025)
ConSep: a Noise- and Reverberation-Robust Speech Separation Framework by Magnitude Conditioning
von: Ho, Kuan-Hsun, et al.
Veröffentlicht: (2024)
von: Ho, Kuan-Hsun, et al.
Veröffentlicht: (2024)
Mitigating Subgroup Disparities in Multi-Label Speech Emotion Recognition: A Pseudo-Labeling and Unsupervised Learning Approach
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2025)
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2025)
What do neural networks listen to? Exploring the crucial bands in Speech Enhancement using Sinc-convolution
von: Ho, Kuan-Hsun, et al.
Veröffentlicht: (2024)
von: Ho, Kuan-Hsun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Enhancing Code-Switching ASR Leveraging Non-Peaky CTC Loss and Deep Language Posterior Injection
von: Yang, Tzu-Ting, et al.
Veröffentlicht: (2024) -
An Effective Context-Balanced Adaptation Approach for Long-Tailed Speech Recognition
von: Wang, Yi-Cheng, et al.
Veröffentlicht: (2024) -
CAMEL: Cross-Attention Enhanced Mixture-of-Experts and Language Bias for Code-Switching Speech Recognition
von: Wang, He, et al.
Veröffentlicht: (2024) -
Enhancing Code-Switching Speech Recognition with LID-Based Collaborative Mixture of Experts Model
von: Huang, Hukai, et al.
Veröffentlicht: (2024) -
Boosting Code-Switching ASR with Mixture of Experts Enhanced Speech-Conditioned LLM
von: Zhang, Fengrun, et al.
Veröffentlicht: (2024)