AdaCS: Adaptive Normalization for Enhanced Code-Switching ASR
Fuente:
arXiv
Guardado en:
| Autores principales: | Chu, The Chuong, Pham, Vu Tuan Dat, Dao, Kien, Nguyen, Hoang, Truong, Quoc Hung |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AsyncSwitch: Asynchronous Text-Speech Adaptation for Code-Switched ASR
por: Nguyen, Tuan, et al.
Publicado: (2025)
por: Nguyen, Tuan, et al.
Publicado: (2025)
Acoustic scattering AI for non-invasive object classifications: A case study on hair assessment
por: Hoang, Long-Vu, et al.
Publicado: (2025)
por: Hoang, Long-Vu, et al.
Publicado: (2025)
Qwen vs. Gemma Integration with Whisper: A Comparative Study in Multilingual SpeechLLM Systems
por: Nguyen, Tuan, et al.
Publicado: (2025)
por: Nguyen, Tuan, et al.
Publicado: (2025)
Pushing the Performance of Synthetic Speech Detection with Kolmogorov-Arnold Networks and Self-Supervised Learning Models
por: Phuong, Tuan Dat, et al.
Publicado: (2025)
por: Phuong, Tuan Dat, et al.
Publicado: (2025)
Can we train ASR systems on Code-switch without real code-switch data? Case study for Singapore's languages
por: Nguyen, Tuan, et al.
Publicado: (2025)
por: Nguyen, Tuan, et al.
Publicado: (2025)
Zero-Shot Text-to-Speech for Vietnamese
por: Vu, Thi, et al.
Publicado: (2025)
por: Vu, Thi, et al.
Publicado: (2025)
VoxVietnam: a Large-Scale Multi-Genre Dataset for Vietnamese Speaker Recognition
por: Vu, Hoang Long, et al.
Publicado: (2024)
por: Vu, Hoang Long, et al.
Publicado: (2024)
Improving Code Switching with Supervised Fine Tuning and GELU Adapters
por: Pham, Linh
Publicado: (2025)
por: Pham, Linh
Publicado: (2025)
Synthetic Data Domain Adaptation for ASR via LLM-based Text and Phonetic Respelling Augmentation
por: Yamashita, Natsuo, et al.
Publicado: (2026)
por: Yamashita, Natsuo, et al.
Publicado: (2026)
DOTA-ME-CS: Daily Oriented Text Audio-Mandarin English-Code Switching Dataset
por: Li, Yupei, et al.
Publicado: (2025)
por: Li, Yupei, et al.
Publicado: (2025)
Boosting Code-Switching ASR with Mixture of Experts Enhanced Speech-Conditioned LLM
por: Zhang, Fengrun, et al.
Publicado: (2024)
por: Zhang, Fengrun, et al.
Publicado: (2024)
Continuous Learning of Transformer-based Audio Deepfake Detection
por: Le, Tuan Duy Nguyen, et al.
Publicado: (2024)
por: Le, Tuan Duy Nguyen, et al.
Publicado: (2024)
CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset
por: Yan, Brian, et al.
Publicado: (2025)
por: Yan, Brian, et al.
Publicado: (2025)
Model-free Speculative Decoding for Transformer-based ASR with Token Map Drafting
por: Ho, Tuan Vu, et al.
Publicado: (2025)
por: Ho, Tuan Vu, et al.
Publicado: (2025)
XLSR-Kanformer: A KAN-Intergrated model for Synthetic Speech Detection
por: Dat, Phuong Tuan, et al.
Publicado: (2025)
por: Dat, Phuong Tuan, et al.
Publicado: (2025)
A Comprehensive Survey with Critical Analysis for Deepfake Speech Detection
por: Pham, Lam, et al.
Publicado: (2024)
por: Pham, Lam, et al.
Publicado: (2024)
MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model
por: Pham, The Hieu, et al.
Publicado: (2025)
por: Pham, The Hieu, et al.
Publicado: (2025)
The Impact of Frequency Bands on Acoustic Anomaly Detection of Machines using Deep Learning Based Model
por: Nguyen, Tin, et al.
Publicado: (2024)
por: Nguyen, Tin, et al.
Publicado: (2024)
Ichigo: Mixed-Modal Early-Fusion Realtime Voice Assistant
por: Dao, Alan, et al.
Publicado: (2024)
por: Dao, Alan, et al.
Publicado: (2024)
SC-MoE: Switch Conformer Mixture of Experts for Unified Streaming and Non-streaming Code-Switching ASR
por: Ye, Shuaishuai, et al.
Publicado: (2024)
por: Ye, Shuaishuai, et al.
Publicado: (2024)
Leave No Knowledge Behind During Knowledge Distillation: Towards Practical and Effective Knowledge Distillation for Code-Switching ASR Using Realistic Data
por: Tseng, Liang-Hsuan, et al.
Publicado: (2024)
por: Tseng, Liang-Hsuan, et al.
Publicado: (2024)
SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS
por: Nguyen, Tan Dat, et al.
Publicado: (2025)
por: Nguyen, Tan Dat, et al.
Publicado: (2025)
Stream-based Active Learning for Anomalous Sound Detection in Machine Condition Monitoring
por: Ho, Tuan Vu, et al.
Publicado: (2024)
por: Ho, Tuan Vu, et al.
Publicado: (2024)
CS-Dialogue: A 104-Hour Dataset of Spontaneous Mandarin-English Code-Switching Dialogues for Speech Recognition
por: Zhou, Jiaming, et al.
Publicado: (2025)
por: Zhou, Jiaming, et al.
Publicado: (2025)
CAMEL: Cross-Attention Enhanced Mixture-of-Experts and Language Bias for Code-Switching Speech Recognition
por: Wang, He, et al.
Publicado: (2024)
por: Wang, He, et al.
Publicado: (2024)
SAML: Speaker Adaptive Mixture of LoRA Experts for End-to-End ASR
por: Zhao, Qiuming, et al.
Publicado: (2024)
por: Zhao, Qiuming, et al.
Publicado: (2024)
Improving Zero-Shot Chinese-English Code-Switching ASR with kNN-CTC and Gated Monolingual Datastores
por: Zhou, Jiaming, et al.
Publicado: (2024)
por: Zhou, Jiaming, et al.
Publicado: (2024)
Attention-Guided Adaptation for Code-Switching Speech Recognition
por: Aditya, Bobbi, et al.
Publicado: (2023)
por: Aditya, Bobbi, et al.
Publicado: (2023)
AdaProj: Adaptively Scaled Angular Margin Subspace Projections for Anomalous Sound Detection with Auxiliary Classification Tasks
por: Wilkinghoff, Kevin
Publicado: (2024)
por: Wilkinghoff, Kevin
Publicado: (2024)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
por: Nguyen, Thai-Binh, et al.
Publicado: (2024)
por: Nguyen, Thai-Binh, et al.
Publicado: (2024)
Overlap-Adaptive Hybrid Speaker Diarization and ASR-Aware Observation Addition for MISP 2025 Challenge
por: Huang, Shangkun, et al.
Publicado: (2025)
por: Huang, Shangkun, et al.
Publicado: (2025)
kNN-CTC: Enhancing ASR via Retrieval of CTC Pseudo Labels
por: Zhou, Jiaming, et al.
Publicado: (2023)
por: Zhou, Jiaming, et al.
Publicado: (2023)
SegAug: CTC-Aligned Segmented Augmentation For Robust RNN-Transducer Based Speech Recognition
por: Le, Khanh, et al.
Publicado: (2025)
por: Le, Khanh, et al.
Publicado: (2025)
ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription
por: Le, Khanh, et al.
Publicado: (2025)
por: Le, Khanh, et al.
Publicado: (2025)
Deepfake Audio Detection Using Spectrogram-based Feature and Ensemble of Deep Learning Models
por: Pham, Lam, et al.
Publicado: (2024)
por: Pham, Lam, et al.
Publicado: (2024)
Enhancing Code-Switching ASR Leveraging Non-Peaky CTC Loss and Deep Language Posterior Injection
por: Yang, Tzu-Ting, et al.
Publicado: (2024)
por: Yang, Tzu-Ting, et al.
Publicado: (2024)
Accent conversion using discrete units with parallel data synthesized from controllable accented TTS
por: Nguyen, Tuan Nam, et al.
Publicado: (2024)
por: Nguyen, Tuan Nam, et al.
Publicado: (2024)
Target Speaker ASR with Whisper
por: Polok, Alexander, et al.
Publicado: (2024)
por: Polok, Alexander, et al.
Publicado: (2024)
Index-ASR Technical Report
por: Song, Zheshu, et al.
Publicado: (2025)
por: Song, Zheshu, et al.
Publicado: (2025)
Room Impulse Responses help attackers to evade Deep Fake Detection
por: Luong, Hieu-Thi, et al.
Publicado: (2024)
por: Luong, Hieu-Thi, et al.
Publicado: (2024)
Ejemplares similares
-
AsyncSwitch: Asynchronous Text-Speech Adaptation for Code-Switched ASR
por: Nguyen, Tuan, et al.
Publicado: (2025) -
Acoustic scattering AI for non-invasive object classifications: A case study on hair assessment
por: Hoang, Long-Vu, et al.
Publicado: (2025) -
Qwen vs. Gemma Integration with Whisper: A Comparative Study in Multilingual SpeechLLM Systems
por: Nguyen, Tuan, et al.
Publicado: (2025) -
Pushing the Performance of Synthetic Speech Detection with Kolmogorov-Arnold Networks and Self-Supervised Learning Models
por: Phuong, Tuan Dat, et al.
Publicado: (2025) -
Can we train ASR systems on Code-switch without real code-switch data? Case study for Singapore's languages
por: Nguyen, Tuan, et al.
Publicado: (2025)