Improving Code Switching with Supervised Fine Tuning and GELU Adapters
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Pham, Linh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Attention-Guided Adaptation for Code-Switching Speech Recognition
von: Aditya, Bobbi, et al.
Veröffentlicht: (2023)
von: Aditya, Bobbi, et al.
Veröffentlicht: (2023)
Improving Audio Captioning Models with Fine-grained Audio Features, Text Embedding Supervision, and LLM Mix-up Augmentation
von: Wu, Shih-Lun, et al.
Veröffentlicht: (2023)
von: Wu, Shih-Lun, et al.
Veröffentlicht: (2023)
Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation
von: Guo, Haohan, et al.
Veröffentlicht: (2024)
von: Guo, Haohan, et al.
Veröffentlicht: (2024)
Fine-Tuning ASR for Stuttered Speech: Personalized vs. Generalized Approaches
von: Mujtaba, Dena, et al.
Veröffentlicht: (2025)
von: Mujtaba, Dena, et al.
Veröffentlicht: (2025)
DOTA-ME-CS: Daily Oriented Text Audio-Mandarin English-Code Switching Dataset
von: Li, Yupei, et al.
Veröffentlicht: (2025)
von: Li, Yupei, et al.
Veröffentlicht: (2025)
CAMEL: Cross-Attention Enhanced Mixture-of-Experts and Language Bias for Code-Switching Speech Recognition
von: Wang, He, et al.
Veröffentlicht: (2024)
von: Wang, He, et al.
Veröffentlicht: (2024)
HingeNet: A Harmonic-Aware Fine-Tuning Approach for Beat Tracking
von: Ru, Ganghui, et al.
Veröffentlicht: (2025)
von: Ru, Ganghui, et al.
Veröffentlicht: (2025)
Efficient Adapter Tuning for Joint Singing Voice Beat and Downbeat Tracking with Self-supervised Learning Features
von: Deng, Jiajun, et al.
Veröffentlicht: (2025)
von: Deng, Jiajun, et al.
Veröffentlicht: (2025)
Adapter Incremental Continual Learning of Efficient Audio Spectrogram Transformers
von: Selvaraj, Nithish Muthuchamy, et al.
Veröffentlicht: (2023)
von: Selvaraj, Nithish Muthuchamy, et al.
Veröffentlicht: (2023)
MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model
von: Pham, The Hieu, et al.
Veröffentlicht: (2025)
von: Pham, The Hieu, et al.
Veröffentlicht: (2025)
Rare Word Recognition and Translation Without Fine-Tuning via Task Vector in Speech Models
von: Jing, Ruihao, et al.
Veröffentlicht: (2025)
von: Jing, Ruihao, et al.
Veröffentlicht: (2025)
Enhancing Polyglot Voices by Leveraging Cross-Lingual Fine-Tuning in Any-to-One Voice Conversion
von: Ruggiero, Giuseppe, et al.
Veröffentlicht: (2024)
von: Ruggiero, Giuseppe, et al.
Veröffentlicht: (2024)
Fine-grained Preference Optimization Improves Zero-shot Text-to-Speech
von: Yao, Jixun, et al.
Veröffentlicht: (2025)
von: Yao, Jixun, et al.
Veröffentlicht: (2025)
Monaural speech enhancement on drone via Adapter based transfer learning
von: Chen, Xingyu, et al.
Veröffentlicht: (2024)
von: Chen, Xingyu, et al.
Veröffentlicht: (2024)
SE/BN Adapter: Parametric Efficient Domain Adaptation for Speaker Recognition
von: Wang, Tianhao, et al.
Veröffentlicht: (2024)
von: Wang, Tianhao, et al.
Veröffentlicht: (2024)
Improving Speech Enhancement with Multi-Metric Supervision from Learned Quality Assessment
von: Wang, Wei, et al.
Veröffentlicht: (2025)
von: Wang, Wei, et al.
Veröffentlicht: (2025)
Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech Recognition
von: Wagner, Dominik, et al.
Veröffentlicht: (2025)
von: Wagner, Dominik, et al.
Veröffentlicht: (2025)
Efficient Emotion and Speaker Adaptation in LLM-Based TTS via Characteristic-Specific Partial Fine-Tuning
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
Omni-R1: Do You Really Need Audio to Fine-Tune Your Audio LLM?
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2025)
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2025)
Speech Emotion Recognition Using Fine-Tuned DWFormer:A Study on Track 1 of the IERPChallenge 2024
von: Wang, Honghong, et al.
Veröffentlicht: (2025)
von: Wang, Honghong, et al.
Veröffentlicht: (2025)
Fine-Tuning Automatic Speech Recognition for People with Parkinson's: An Effective Strategy for Enhancing Speech Technology Accessibility
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2024)
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2024)
VoiceTailor: Lightweight Plug-In Adapter for Diffusion-Based Personalized Text-to-Speech
von: Kim, Heeseung, et al.
Veröffentlicht: (2024)
von: Kim, Heeseung, et al.
Veröffentlicht: (2024)
Metadata-Enhanced Speech Emotion Recognition: Augmented Residual Integration and Co-Attention in Two-Stage Fine-Tuning
von: Wan, Zixiang, et al.
Veröffentlicht: (2024)
von: Wan, Zixiang, et al.
Veröffentlicht: (2024)
Efficient Adapter Tuning of Pre-trained Speech Models for Automatic Speaker Verification
von: Sang, Mufan, et al.
Veröffentlicht: (2024)
von: Sang, Mufan, et al.
Veröffentlicht: (2024)
HDMoLE: Mixture of LoRA Experts with Hierarchical Routing and Dynamic Thresholds for Fine-Tuning LLM-based ASR Models
von: Mu, Bingshen, et al.
Veröffentlicht: (2024)
von: Mu, Bingshen, et al.
Veröffentlicht: (2024)
DQLoRA: A Lightweight Domain-Aware Denoising ASR via Adapter-guided Distillation
von: Yang, Yiru
Veröffentlicht: (2025)
von: Yang, Yiru
Veröffentlicht: (2025)
Adapting Language Balance in Code-Switching Speech
von: Ugan, Enes Yavuz, et al.
Veröffentlicht: (2025)
von: Ugan, Enes Yavuz, et al.
Veröffentlicht: (2025)
AsyncSwitch: Asynchronous Text-Speech Adaptation for Code-Switched ASR
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025)
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025)
Windowed SummaryMixing: An Efficient Fine-Tuning of Self-Supervised Learning Models for Low-resource Speech Recognition
von: Menon, Aditya Srinivas, et al.
Veröffentlicht: (2026)
von: Menon, Aditya Srinivas, et al.
Veröffentlicht: (2026)
T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback
von: Wang, Zehan, et al.
Veröffentlicht: (2025)
von: Wang, Zehan, et al.
Veröffentlicht: (2025)
Crab: Multi Layer Contrastive Supervision to Improve Speech Emotion Recognition Under Both Acted and Natural Speech Condition
von: Ueda, Lucas H., et al.
Veröffentlicht: (2026)
von: Ueda, Lucas H., et al.
Veröffentlicht: (2026)
Improving Zero-Shot Chinese-English Code-Switching ASR with kNN-CTC and Gated Monolingual Datastores
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
Jointly Fine-Tuning "BERT-like" Self Supervised Models to Improve Multimodal Speech Emotion Recognition
von: Siriwardhana, Shamane, et al.
Veröffentlicht: (2020)
von: Siriwardhana, Shamane, et al.
Veröffentlicht: (2020)
Improving Anomalous Sound Detection via Low-Rank Adaptation Fine-Tuning of Pre-Trained Audio Models
von: Zheng, Xinhu, et al.
Veröffentlicht: (2024)
von: Zheng, Xinhu, et al.
Veröffentlicht: (2024)
SwitchCodec: A High-Fidelity Nerual Audio Codec With Sparse Quantization
von: Wang, Jin, et al.
Veröffentlicht: (2025)
von: Wang, Jin, et al.
Veröffentlicht: (2025)
SC-MoE: Switch Conformer Mixture of Experts for Unified Streaming and Non-streaming Code-Switching ASR
von: Ye, Shuaishuai, et al.
Veröffentlicht: (2024)
von: Ye, Shuaishuai, et al.
Veröffentlicht: (2024)
Xi+: Uncertainty Supervision for Robust Speaker Embedding
von: Li, Junjie, et al.
Veröffentlicht: (2025)
von: Li, Junjie, et al.
Veröffentlicht: (2025)
Contextual Biasing to Improve Domain-specific Custom Vocabulary Audio Transcription without Explicit Fine-Tuning of Whisper Model
von: Lall, Vishakha, et al.
Veröffentlicht: (2024)
von: Lall, Vishakha, et al.
Veröffentlicht: (2024)
Enhanced Reverberation as Supervision for Unsupervised Speech Separation
von: Saijo, Kohei, et al.
Veröffentlicht: (2024)
von: Saijo, Kohei, et al.
Veröffentlicht: (2024)
Leveraging Self-Supervised Learning for Speaker Diarization
von: Han, Jiangyu, et al.
Veröffentlicht: (2024)
von: Han, Jiangyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Attention-Guided Adaptation for Code-Switching Speech Recognition
von: Aditya, Bobbi, et al.
Veröffentlicht: (2023) -
Improving Audio Captioning Models with Fine-grained Audio Features, Text Embedding Supervision, and LLM Mix-up Augmentation
von: Wu, Shih-Lun, et al.
Veröffentlicht: (2023) -
Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation
von: Guo, Haohan, et al.
Veröffentlicht: (2024) -
Fine-Tuning ASR for Stuttered Speech: Personalized vs. Generalized Approaches
von: Mujtaba, Dena, et al.
Veröffentlicht: (2025) -
DOTA-ME-CS: Daily Oriented Text Audio-Mandarin English-Code Switching Dataset
von: Li, Yupei, et al.
Veröffentlicht: (2025)