SC-MoE: Switch Conformer Mixture of Experts for Unified Streaming and Non-streaming Code-Switching ASR
Fuente:
arXiv
Guardado en:
| Autores principales: | Ye, Shuaishuai, Chen, Shunfei, Hu, Xinhui, Xu, Xinkang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The RoyalFlush Automatic Speech Diarization and Recognition System for In-Car Multi-Channel Automatic Speech Recognition Challenge
por: Tian, Jingguang, et al.
Publicado: (2024)
por: Tian, Jingguang, et al.
Publicado: (2024)
Boosting Code-Switching ASR with Mixture of Experts Enhanced Speech-Conditioned LLM
por: Zhang, Fengrun, et al.
Publicado: (2024)
por: Zhang, Fengrun, et al.
Publicado: (2024)
Learning Emotion-Invariant Speaker Representations for Speaker Verification
por: Tian, Jingguang, et al.
Publicado: (2025)
por: Tian, Jingguang, et al.
Publicado: (2025)
Unifying Streaming and Non-streaming Zipformer-based ASR
por: Sharma, Bidisha, et al.
Publicado: (2025)
por: Sharma, Bidisha, et al.
Publicado: (2025)
Discrete Audio Representations for Automated Audio Captioning
por: Tian, Jingguang, et al.
Publicado: (2025)
por: Tian, Jingguang, et al.
Publicado: (2025)
BLR-MoE: Boosted Language-Routing Mixture of Experts for Domain-Robust Multilingual E2E ASR
por: Ma, Guodong, et al.
Publicado: (2025)
por: Ma, Guodong, et al.
Publicado: (2025)
CAMEL: Cross-Attention Enhanced Mixture-of-Experts and Language Bias for Code-Switching Speech Recognition
por: Wang, He, et al.
Publicado: (2024)
por: Wang, He, et al.
Publicado: (2024)
AsyncSwitch: Asynchronous Text-Speech Adaptation for Code-Switched ASR
por: Nguyen, Tuan, et al.
Publicado: (2025)
por: Nguyen, Tuan, et al.
Publicado: (2025)
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
por: Xue, Heyang, et al.
Publicado: (2025)
por: Xue, Heyang, et al.
Publicado: (2025)
Enhancing Code-Switching Speech Recognition with LID-Based Collaborative Mixture of Experts Model
por: Huang, Hukai, et al.
Publicado: (2024)
por: Huang, Hukai, et al.
Publicado: (2024)
SAML: Speaker Adaptive Mixture of LoRA Experts for End-to-End ASR
por: Zhao, Qiuming, et al.
Publicado: (2024)
por: Zhao, Qiuming, et al.
Publicado: (2024)
Attention-Guided Adaptation for Code-Switching Speech Recognition
por: Aditya, Bobbi, et al.
Publicado: (2023)
por: Aditya, Bobbi, et al.
Publicado: (2023)
A Language-Agnostic Hierarchical LoRA-MoE Architecture for CTC-based Multilingual ASR
por: Zheng, Yuang, et al.
Publicado: (2026)
por: Zheng, Yuang, et al.
Publicado: (2026)
HDMoLE: Mixture of LoRA Experts with Hierarchical Routing and Dynamic Thresholds for Fine-Tuning LLM-based ASR Models
por: Mu, Bingshen, et al.
Publicado: (2024)
por: Mu, Bingshen, et al.
Publicado: (2024)
Improving Code Switching with Supervised Fine Tuning and GELU Adapters
por: Pham, Linh
Publicado: (2025)
por: Pham, Linh
Publicado: (2025)
Decoder-only Conformer with Modality-aware Sparse Mixtures of Experts for ASR
por: Lee, Jaeyoung, et al.
Publicado: (2026)
por: Lee, Jaeyoung, et al.
Publicado: (2026)
An Effective Mixture-Of-Experts Approach For Code-Switching Speech Recognition Leveraging Encoder Disentanglement
por: Yang, Tzu-Ting, et al.
Publicado: (2024)
por: Yang, Tzu-Ting, et al.
Publicado: (2024)
AdaCS: Adaptive Normalization for Enhanced Code-Switching ASR
por: Chu, The Chuong, et al.
Publicado: (2025)
por: Chu, The Chuong, et al.
Publicado: (2025)
RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
por: Sun, Haoqin, et al.
Publicado: (2025)
por: Sun, Haoqin, et al.
Publicado: (2025)
Improving Zero-Shot Chinese-English Code-Switching ASR with kNN-CTC and Gated Monolingual Datastores
por: Zhou, Jiaming, et al.
Publicado: (2024)
por: Zhou, Jiaming, et al.
Publicado: (2024)
Enhancing Code-Switching ASR Leveraging Non-Peaky CTC Loss and Deep Language Posterior Injection
por: Yang, Tzu-Ting, et al.
Publicado: (2024)
por: Yang, Tzu-Ting, et al.
Publicado: (2024)
Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams
por: He, Xiluo, et al.
Publicado: (2025)
por: He, Xiluo, et al.
Publicado: (2025)
DOTA-ME-CS: Daily Oriented Text Audio-Mandarin English-Code Switching Dataset
por: Li, Yupei, et al.
Publicado: (2025)
por: Li, Yupei, et al.
Publicado: (2025)
Towards One-bit ASR: Extremely Low-bit Conformer Quantization Using Co-training and Stochastic Precision
por: Li, Zhaoqing, et al.
Publicado: (2025)
por: Li, Zhaoqing, et al.
Publicado: (2025)
DualVC 2: Dynamic Masked Convolution for Unified Streaming and Non-Streaming Voice Conversion
por: Ning, Ziqian, et al.
Publicado: (2023)
por: Ning, Ziqian, et al.
Publicado: (2023)
Lightweight Target-Speaker-Based Overlap Transcription for Practical Streaming ASR
por: Pražák, Aleš, et al.
Publicado: (2025)
por: Pražák, Aleš, et al.
Publicado: (2025)
CUSIDE-T: Chunking, Simulating Future and Decoding for Transducer based Streaming ASR
por: Zhao, Wenbo, et al.
Publicado: (2024)
por: Zhao, Wenbo, et al.
Publicado: (2024)
SSCFormer: Push the Limit of Chunk-wise Conformer for Streaming ASR Using Sequentially Sampled Chunks and Chunked Causal Convolution
por: Wang, Fangyuan, et al.
Publicado: (2022)
por: Wang, Fangyuan, et al.
Publicado: (2022)
DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec
por: Chen, Peijie, et al.
Publicado: (2025)
por: Chen, Peijie, et al.
Publicado: (2025)
Leave No Knowledge Behind During Knowledge Distillation: Towards Practical and Effective Knowledge Distillation for Code-Switching ASR Using Realistic Data
por: Tseng, Liang-Hsuan, et al.
Publicado: (2024)
por: Tseng, Liang-Hsuan, et al.
Publicado: (2024)
On-the-fly Routing for Zero-shot MoE Speaker Adaptation of Speech Foundation Models for Dysarthric Speech Recognition
por: HU, Shujie, et al.
Publicado: (2025)
por: HU, Shujie, et al.
Publicado: (2025)
DCTX-Conformer: Dynamic context carry-over for low latency unified streaming and non-streaming Conformer ASR
por: Huybrechts, Goeric, et al.
Publicado: (2023)
por: Huybrechts, Goeric, et al.
Publicado: (2023)
Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency Streaming ASR
por: Li, Longhao, et al.
Publicado: (2025)
por: Li, Longhao, et al.
Publicado: (2025)
Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware Decoding
por: Zhao, Jiahui, et al.
Publicado: (2024)
por: Zhao, Jiahui, et al.
Publicado: (2024)
A Comparative Study of LLM-based ASR and Whisper in Low Resource and Code Switching Scenario
por: Song, Zheshu, et al.
Publicado: (2024)
por: Song, Zheshu, et al.
Publicado: (2024)
A Unified Denoising and Adaptation Framework for Self-Supervised Bengali Dialectal ASR
por: Biswas, Swadhin, et al.
Publicado: (2025)
por: Biswas, Swadhin, et al.
Publicado: (2025)
DiaMoE-TTS: A Unified IPA-Based Dialect TTS Framework with Mixture-of-Experts and Parameter-Efficient Zero-Shot Adaptation
por: Chen, Ziqi, et al.
Publicado: (2025)
por: Chen, Ziqi, et al.
Publicado: (2025)
Promptformer: Prompted Conformer Transducer for ASR
por: Duarte-Torres, Sergio, et al.
Publicado: (2024)
por: Duarte-Torres, Sergio, et al.
Publicado: (2024)
Unraveling Complex Data Diversity in Underwater Acoustic Target Recognition through Convolution-based Mixture of Experts
por: Xie, Yuan, et al.
Publicado: (2024)
por: Xie, Yuan, et al.
Publicado: (2024)
SwitchCodec: A High-Fidelity Nerual Audio Codec With Sparse Quantization
por: Wang, Jin, et al.
Publicado: (2025)
por: Wang, Jin, et al.
Publicado: (2025)
Ejemplares similares
-
The RoyalFlush Automatic Speech Diarization and Recognition System for In-Car Multi-Channel Automatic Speech Recognition Challenge
por: Tian, Jingguang, et al.
Publicado: (2024) -
Boosting Code-Switching ASR with Mixture of Experts Enhanced Speech-Conditioned LLM
por: Zhang, Fengrun, et al.
Publicado: (2024) -
Learning Emotion-Invariant Speaker Representations for Speaker Verification
por: Tian, Jingguang, et al.
Publicado: (2025) -
Unifying Streaming and Non-streaming Zipformer-based ASR
por: Sharma, Bidisha, et al.
Publicado: (2025) -
Discrete Audio Representations for Automated Audio Captioning
por: Tian, Jingguang, et al.
Publicado: (2025)