Optimizing ASR for Catalan-Spanish Code-Switching: A Comparative Analysis of Methodologies
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mena, Carlos, Serra, Pol, Romero, Jacobo, Messaoudi, Abir, Giraldo, Jose, Armentano-Oller, Carme, Zevallos, Rodolfo, Meza, Ivan, Hernando, Javier |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enhancing Crowdsourced Audio for Text-to-Speech Models
von: Giraldo, José, et al.
Veröffentlicht: (2024)
von: Giraldo, José, et al.
Veröffentlicht: (2024)
How Attention Shapes Emotion: A Comparative Study of Attention Mechanisms for Speech Emotion Recognition
von: Casals-Salvador, Marc, et al.
Veröffentlicht: (2026)
von: Casals-Salvador, Marc, et al.
Veröffentlicht: (2026)
Zero-Shot TTS With Enhanced Audio Prompts: Bsc Submission For The 2026 Wildspoof Challenge TTS Track
von: Giraldo, Jose, et al.
Veröffentlicht: (2026)
von: Giraldo, Jose, et al.
Veröffentlicht: (2026)
Semi-supervised Learning for Code-Switching ASR with Large Language Model Filter
von: Xi, Yu, et al.
Veröffentlicht: (2024)
von: Xi, Yu, et al.
Veröffentlicht: (2024)
Doctor or Patient? Synergizing Diarization and ASR for Code-Switched Hinglish Medical Conditions Extraction
von: Baroudi, Séverin, et al.
Veröffentlicht: (2026)
von: Baroudi, Séverin, et al.
Veröffentlicht: (2026)
AsyncSwitch: Asynchronous Text-Speech Adaptation for Code-Switched ASR
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025)
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025)
Boosting Code-Switching ASR with Mixture of Experts Enhanced Speech-Conditioned LLM
von: Zhang, Fengrun, et al.
Veröffentlicht: (2024)
von: Zhang, Fengrun, et al.
Veröffentlicht: (2024)
A Comparative Study of LLM-based ASR and Whisper in Low Resource and Code Switching Scenario
von: Song, Zheshu, et al.
Veröffentlicht: (2024)
von: Song, Zheshu, et al.
Veröffentlicht: (2024)
SC-MoE: Switch Conformer Mixture of Experts for Unified Streaming and Non-streaming Code-Switching ASR
von: Ye, Shuaishuai, et al.
Veröffentlicht: (2024)
von: Ye, Shuaishuai, et al.
Veröffentlicht: (2024)
AdaCS: Adaptive Normalization for Enhanced Code-Switching ASR
von: Chu, The Chuong, et al.
Veröffentlicht: (2025)
von: Chu, The Chuong, et al.
Veröffentlicht: (2025)
Comparative Analysis of ASR Methods for Speech Deepfake Detection
von: Salvi, Davide, et al.
Veröffentlicht: (2024)
von: Salvi, Davide, et al.
Veröffentlicht: (2024)
Comparing Unsupervised and Supervised Semantic Speech Tokens: A Case Study of Child ASR
von: Shi, Mohan, et al.
Veröffentlicht: (2025)
von: Shi, Mohan, et al.
Veröffentlicht: (2025)
Improving Zero-Shot Chinese-English Code-Switching ASR with kNN-CTC and Gated Monolingual Datastores
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models
von: Li, Li, et al.
Veröffentlicht: (2026)
von: Li, Li, et al.
Veröffentlicht: (2026)
First Steps Towards Voice Anonymization for Code-Switching Speech
von: Meyer, Sarina, et al.
Veröffentlicht: (2025)
von: Meyer, Sarina, et al.
Veröffentlicht: (2025)
NeuRO: An Application for Code-Switched Autism Detection in Children
von: Akhtar, Mohd Mujtaba, et al.
Veröffentlicht: (2024)
von: Akhtar, Mohd Mujtaba, et al.
Veröffentlicht: (2024)
MDM-ASR: Bridging Accuracy and Efficiency in ASR with Diffusion-Based Non-Autoregressive Decoding
von: Yen, Hao, et al.
Veröffentlicht: (2026)
von: Yen, Hao, et al.
Veröffentlicht: (2026)
Mind the Gap: Impact of Synthetic Conversational Data on Multi-Talker ASR and Speaker Diarization
von: Polok, Alexander, et al.
Veröffentlicht: (2026)
von: Polok, Alexander, et al.
Veröffentlicht: (2026)
Investigating Zero-Shot Generalizability on Mandarin-English Code-Switched ASR and Speech-to-text Translation of Recent Foundation Models with Self-Supervision and Weak Supervision
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2023)
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2023)
All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASR
von: Moriya, Takafumi, et al.
Veröffentlicht: (2025)
von: Moriya, Takafumi, et al.
Veröffentlicht: (2025)
Quantifying Cross-Lingual Transfer in Paralinguistic Speech Tasks
von: Buitrago, Pol, et al.
Veröffentlicht: (2026)
von: Buitrago, Pol, et al.
Veröffentlicht: (2026)
Leave No Knowledge Behind During Knowledge Distillation: Towards Practical and Effective Knowledge Distillation for Code-Switching ASR Using Realistic Data
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2024)
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2024)
Attention-Guided Adaptation for Code-Switching Speech Recognition
von: Aditya, Bobbi, et al.
Veröffentlicht: (2023)
von: Aditya, Bobbi, et al.
Veröffentlicht: (2023)
Inverse-Hessian Regularization for Continual Learning in ASR
von: Eeckt, Steven Vander, et al.
Veröffentlicht: (2026)
von: Eeckt, Steven Vander, et al.
Veröffentlicht: (2026)
Improving Code Switching with Supervised Fine Tuning and GELU Adapters
von: Pham, Linh
Veröffentlicht: (2025)
von: Pham, Linh
Veröffentlicht: (2025)
SOT Triggered Neural Clustering for Speaker Attributed ASR
von: Zheng, Xianrui, et al.
Veröffentlicht: (2024)
von: Zheng, Xianrui, et al.
Veröffentlicht: (2024)
DNCASR: End-to-End Training for Speaker-Attributed ASR
von: Zheng, Xianrui, et al.
Veröffentlicht: (2025)
von: Zheng, Xianrui, et al.
Veröffentlicht: (2025)
An investigation of modularity for noise robustness in conformer-based ASR
von: de Gibson, Louise Coppieters, et al.
Veröffentlicht: (2024)
von: de Gibson, Louise Coppieters, et al.
Veröffentlicht: (2024)
Enhancing Code-Switching ASR Leveraging Non-Peaky CTC Loss and Deep Language Posterior Injection
von: Yang, Tzu-Ting, et al.
Veröffentlicht: (2024)
von: Yang, Tzu-Ting, et al.
Veröffentlicht: (2024)
Resource-Efficient Adaptation of Speech Foundation Models for Multi-Speaker ASR
von: Wang, Weiqing, et al.
Veröffentlicht: (2024)
von: Wang, Weiqing, et al.
Veröffentlicht: (2024)
Speaker Targeting via Self-Speaker Adaptation for Multi-talker ASR
von: Wang, Weiqing, et al.
Veröffentlicht: (2025)
von: Wang, Weiqing, et al.
Veröffentlicht: (2025)
Bootstrapping Audiovisual Speech Recognition in Zero-AV-Resource Scenarios with Synthetic Visual Data
von: Buitrago, Pol, et al.
Veröffentlicht: (2026)
von: Buitrago, Pol, et al.
Veröffentlicht: (2026)
Target Speaker ASR with Whisper
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
Index-ASR Technical Report
von: Song, Zheshu, et al.
Veröffentlicht: (2025)
von: Song, Zheshu, et al.
Veröffentlicht: (2025)
NLE: Non-autoregressive LLM-based ASR by Transcript Editing
von: Dekel, Avihu, et al.
Veröffentlicht: (2026)
von: Dekel, Avihu, et al.
Veröffentlicht: (2026)
XLSR-Transducer: Streaming ASR for Self-Supervised Pretrained Models
von: Kumar, Shashi, et al.
Veröffentlicht: (2024)
von: Kumar, Shashi, et al.
Veröffentlicht: (2024)
Towards a Single ASR Model That Generalizes to Disordered Speech
von: Tobin, Jimmy, et al.
Veröffentlicht: (2024)
von: Tobin, Jimmy, et al.
Veröffentlicht: (2024)
Comparative Evaluation of Text and Audio Simplification: A Methodological Replication Study
von: Barai, Prosanta, et al.
Veröffentlicht: (2025)
von: Barai, Prosanta, et al.
Veröffentlicht: (2025)
MedASR: An Open-Source Model for High-Accuracy Medical Dictation
von: Wu, Ke, et al.
Veröffentlicht: (2026)
von: Wu, Ke, et al.
Veröffentlicht: (2026)
Contextual Biasing for Streaming ASR via CTC-based Word Spotting
von: Tsai, Kai-Chen, et al.
Veröffentlicht: (2026)
von: Tsai, Kai-Chen, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Enhancing Crowdsourced Audio for Text-to-Speech Models
von: Giraldo, José, et al.
Veröffentlicht: (2024) -
How Attention Shapes Emotion: A Comparative Study of Attention Mechanisms for Speech Emotion Recognition
von: Casals-Salvador, Marc, et al.
Veröffentlicht: (2026) -
Zero-Shot TTS With Enhanced Audio Prompts: Bsc Submission For The 2026 Wildspoof Challenge TTS Track
von: Giraldo, Jose, et al.
Veröffentlicht: (2026) -
Semi-supervised Learning for Code-Switching ASR with Large Language Model Filter
von: Xi, Yu, et al.
Veröffentlicht: (2024) -
Doctor or Patient? Synergizing Diarization and ASR for Code-Switched Hinglish Medical Conditions Extraction
von: Baroudi, Séverin, et al.
Veröffentlicht: (2026)