AdaCS: Adaptive Normalization for Enhanced Code-Switching ASR
Fuente:
arXiv
Salvato in:
| Autori principali: | Chu, The Chuong, Pham, Vu Tuan Dat, Dao, Kien, Nguyen, Hoang, Truong, Quoc Hung |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AsyncSwitch: Asynchronous Text-Speech Adaptation for Code-Switched ASR
di: Nguyen, Tuan, et al.
Pubblicazione: (2025)
di: Nguyen, Tuan, et al.
Pubblicazione: (2025)
Acoustic scattering AI for non-invasive object classifications: A case study on hair assessment
di: Hoang, Long-Vu, et al.
Pubblicazione: (2025)
di: Hoang, Long-Vu, et al.
Pubblicazione: (2025)
Qwen vs. Gemma Integration with Whisper: A Comparative Study in Multilingual SpeechLLM Systems
di: Nguyen, Tuan, et al.
Pubblicazione: (2025)
di: Nguyen, Tuan, et al.
Pubblicazione: (2025)
Pushing the Performance of Synthetic Speech Detection with Kolmogorov-Arnold Networks and Self-Supervised Learning Models
di: Phuong, Tuan Dat, et al.
Pubblicazione: (2025)
di: Phuong, Tuan Dat, et al.
Pubblicazione: (2025)
Can we train ASR systems on Code-switch without real code-switch data? Case study for Singapore's languages
di: Nguyen, Tuan, et al.
Pubblicazione: (2025)
di: Nguyen, Tuan, et al.
Pubblicazione: (2025)
Zero-Shot Text-to-Speech for Vietnamese
di: Vu, Thi, et al.
Pubblicazione: (2025)
di: Vu, Thi, et al.
Pubblicazione: (2025)
VoxVietnam: a Large-Scale Multi-Genre Dataset for Vietnamese Speaker Recognition
di: Vu, Hoang Long, et al.
Pubblicazione: (2024)
di: Vu, Hoang Long, et al.
Pubblicazione: (2024)
Improving Code Switching with Supervised Fine Tuning and GELU Adapters
di: Pham, Linh
Pubblicazione: (2025)
di: Pham, Linh
Pubblicazione: (2025)
Synthetic Data Domain Adaptation for ASR via LLM-based Text and Phonetic Respelling Augmentation
di: Yamashita, Natsuo, et al.
Pubblicazione: (2026)
di: Yamashita, Natsuo, et al.
Pubblicazione: (2026)
DOTA-ME-CS: Daily Oriented Text Audio-Mandarin English-Code Switching Dataset
di: Li, Yupei, et al.
Pubblicazione: (2025)
di: Li, Yupei, et al.
Pubblicazione: (2025)
Boosting Code-Switching ASR with Mixture of Experts Enhanced Speech-Conditioned LLM
di: Zhang, Fengrun, et al.
Pubblicazione: (2024)
di: Zhang, Fengrun, et al.
Pubblicazione: (2024)
Continuous Learning of Transformer-based Audio Deepfake Detection
di: Le, Tuan Duy Nguyen, et al.
Pubblicazione: (2024)
di: Le, Tuan Duy Nguyen, et al.
Pubblicazione: (2024)
CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset
di: Yan, Brian, et al.
Pubblicazione: (2025)
di: Yan, Brian, et al.
Pubblicazione: (2025)
Model-free Speculative Decoding for Transformer-based ASR with Token Map Drafting
di: Ho, Tuan Vu, et al.
Pubblicazione: (2025)
di: Ho, Tuan Vu, et al.
Pubblicazione: (2025)
XLSR-Kanformer: A KAN-Intergrated model for Synthetic Speech Detection
di: Dat, Phuong Tuan, et al.
Pubblicazione: (2025)
di: Dat, Phuong Tuan, et al.
Pubblicazione: (2025)
A Comprehensive Survey with Critical Analysis for Deepfake Speech Detection
di: Pham, Lam, et al.
Pubblicazione: (2024)
di: Pham, Lam, et al.
Pubblicazione: (2024)
MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model
di: Pham, The Hieu, et al.
Pubblicazione: (2025)
di: Pham, The Hieu, et al.
Pubblicazione: (2025)
The Impact of Frequency Bands on Acoustic Anomaly Detection of Machines using Deep Learning Based Model
di: Nguyen, Tin, et al.
Pubblicazione: (2024)
di: Nguyen, Tin, et al.
Pubblicazione: (2024)
Ichigo: Mixed-Modal Early-Fusion Realtime Voice Assistant
di: Dao, Alan, et al.
Pubblicazione: (2024)
di: Dao, Alan, et al.
Pubblicazione: (2024)
SC-MoE: Switch Conformer Mixture of Experts for Unified Streaming and Non-streaming Code-Switching ASR
di: Ye, Shuaishuai, et al.
Pubblicazione: (2024)
di: Ye, Shuaishuai, et al.
Pubblicazione: (2024)
Leave No Knowledge Behind During Knowledge Distillation: Towards Practical and Effective Knowledge Distillation for Code-Switching ASR Using Realistic Data
di: Tseng, Liang-Hsuan, et al.
Pubblicazione: (2024)
di: Tseng, Liang-Hsuan, et al.
Pubblicazione: (2024)
SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS
di: Nguyen, Tan Dat, et al.
Pubblicazione: (2025)
di: Nguyen, Tan Dat, et al.
Pubblicazione: (2025)
Stream-based Active Learning for Anomalous Sound Detection in Machine Condition Monitoring
di: Ho, Tuan Vu, et al.
Pubblicazione: (2024)
di: Ho, Tuan Vu, et al.
Pubblicazione: (2024)
CS-Dialogue: A 104-Hour Dataset of Spontaneous Mandarin-English Code-Switching Dialogues for Speech Recognition
di: Zhou, Jiaming, et al.
Pubblicazione: (2025)
di: Zhou, Jiaming, et al.
Pubblicazione: (2025)
CAMEL: Cross-Attention Enhanced Mixture-of-Experts and Language Bias for Code-Switching Speech Recognition
di: Wang, He, et al.
Pubblicazione: (2024)
di: Wang, He, et al.
Pubblicazione: (2024)
SAML: Speaker Adaptive Mixture of LoRA Experts for End-to-End ASR
di: Zhao, Qiuming, et al.
Pubblicazione: (2024)
di: Zhao, Qiuming, et al.
Pubblicazione: (2024)
Improving Zero-Shot Chinese-English Code-Switching ASR with kNN-CTC and Gated Monolingual Datastores
di: Zhou, Jiaming, et al.
Pubblicazione: (2024)
di: Zhou, Jiaming, et al.
Pubblicazione: (2024)
Attention-Guided Adaptation for Code-Switching Speech Recognition
di: Aditya, Bobbi, et al.
Pubblicazione: (2023)
di: Aditya, Bobbi, et al.
Pubblicazione: (2023)
AdaProj: Adaptively Scaled Angular Margin Subspace Projections for Anomalous Sound Detection with Auxiliary Classification Tasks
di: Wilkinghoff, Kevin
Pubblicazione: (2024)
di: Wilkinghoff, Kevin
Pubblicazione: (2024)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
di: Nguyen, Thai-Binh, et al.
Pubblicazione: (2024)
di: Nguyen, Thai-Binh, et al.
Pubblicazione: (2024)
Overlap-Adaptive Hybrid Speaker Diarization and ASR-Aware Observation Addition for MISP 2025 Challenge
di: Huang, Shangkun, et al.
Pubblicazione: (2025)
di: Huang, Shangkun, et al.
Pubblicazione: (2025)
kNN-CTC: Enhancing ASR via Retrieval of CTC Pseudo Labels
di: Zhou, Jiaming, et al.
Pubblicazione: (2023)
di: Zhou, Jiaming, et al.
Pubblicazione: (2023)
SegAug: CTC-Aligned Segmented Augmentation For Robust RNN-Transducer Based Speech Recognition
di: Le, Khanh, et al.
Pubblicazione: (2025)
di: Le, Khanh, et al.
Pubblicazione: (2025)
ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription
di: Le, Khanh, et al.
Pubblicazione: (2025)
di: Le, Khanh, et al.
Pubblicazione: (2025)
Deepfake Audio Detection Using Spectrogram-based Feature and Ensemble of Deep Learning Models
di: Pham, Lam, et al.
Pubblicazione: (2024)
di: Pham, Lam, et al.
Pubblicazione: (2024)
Enhancing Code-Switching ASR Leveraging Non-Peaky CTC Loss and Deep Language Posterior Injection
di: Yang, Tzu-Ting, et al.
Pubblicazione: (2024)
di: Yang, Tzu-Ting, et al.
Pubblicazione: (2024)
Accent conversion using discrete units with parallel data synthesized from controllable accented TTS
di: Nguyen, Tuan Nam, et al.
Pubblicazione: (2024)
di: Nguyen, Tuan Nam, et al.
Pubblicazione: (2024)
Target Speaker ASR with Whisper
di: Polok, Alexander, et al.
Pubblicazione: (2024)
di: Polok, Alexander, et al.
Pubblicazione: (2024)
Index-ASR Technical Report
di: Song, Zheshu, et al.
Pubblicazione: (2025)
di: Song, Zheshu, et al.
Pubblicazione: (2025)
Room Impulse Responses help attackers to evade Deep Fake Detection
di: Luong, Hieu-Thi, et al.
Pubblicazione: (2024)
di: Luong, Hieu-Thi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
AsyncSwitch: Asynchronous Text-Speech Adaptation for Code-Switched ASR
di: Nguyen, Tuan, et al.
Pubblicazione: (2025) -
Acoustic scattering AI for non-invasive object classifications: A case study on hair assessment
di: Hoang, Long-Vu, et al.
Pubblicazione: (2025) -
Qwen vs. Gemma Integration with Whisper: A Comparative Study in Multilingual SpeechLLM Systems
di: Nguyen, Tuan, et al.
Pubblicazione: (2025) -
Pushing the Performance of Synthetic Speech Detection with Kolmogorov-Arnold Networks and Self-Supervised Learning Models
di: Phuong, Tuan Dat, et al.
Pubblicazione: (2025) -
Can we train ASR systems on Code-switch without real code-switch data? Case study for Singapore's languages
di: Nguyen, Tuan, et al.
Pubblicazione: (2025)