How to Learn a New Language? An Efficient Solution for Self-Supervised Learning Models Unseen Languages Adaption in Low-Resource Scenario
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wang, Shih-Heng, Chen, Zih-Ching, Shi, Jiatong, Chuang, Ming-To, Lin, Guan-Ting, Huang, Kuan-Po, Harwath, David, Li, Shang-Wen, Lee, Hung-yi |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages
par: Huang, Kuan-Po, et autres
Publié: (2023)
par: Huang, Kuan-Po, et autres
Publié: (2023)
Low-Resource Self-Supervised Learning with SSL-Enhanced TTS
par: Hsu, Po-chun, et autres
Publié: (2023)
par: Hsu, Po-chun, et autres
Publié: (2023)
Do Neural Codecs Generalize? A Controlled Study Across Unseen Languages and Non-Speech Tasks
par: Wang, Shih-Heng, et autres
Publié: (2026)
par: Wang, Shih-Heng, et autres
Publié: (2026)
Interface Design for Self-Supervised Speech Models
par: Shih, Yi-Jen, et autres
Publié: (2024)
par: Shih, Yi-Jen, et autres
Publié: (2024)
SMILE: Speech Meta In-Context Learning for Low-Resource Language Automatic Speech Recognition
par: Hsu, Ming-Hao, et autres
Publié: (2024)
par: Hsu, Ming-Hao, et autres
Publié: (2024)
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
par: Wang, Shih-heng, et autres
Publié: (2024)
par: Wang, Shih-heng, et autres
Publié: (2024)
EFFUSE: Efficient Self-Supervised Feature Fusion for E2E ASR in Low Resource and Multilingual Scenarios
par: Srivastava, Tejes, et autres
Publié: (2023)
par: Srivastava, Tejes, et autres
Publié: (2023)
Can Large Audio-Language Models Truly Hear? Tackling Hallucinations with Multi-Task Assessment and Stepwise Audio Reasoning
par: Kuan, Chun-Yi, et autres
Publié: (2024)
par: Kuan, Chun-Yi, et autres
Publié: (2024)
Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples
par: Kuan, Chun-Yi, et autres
Publié: (2025)
par: Kuan, Chun-Yi, et autres
Publié: (2025)
Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy Speech
par: Lin, Guan-Ting, et autres
Publié: (2024)
par: Lin, Guan-Ting, et autres
Publié: (2024)
Do Prompts Really Prompt? Exploring the Prompt Understanding Capability of Whisper
par: Yang, Chih-Kai, et autres
Publié: (2024)
par: Yang, Chih-Kai, et autres
Publié: (2024)
ASPIRin: Action Space Projection for Interactivity-Optimized Reinforcement Learning in Full-Duplex Speech Language Models
par: Hsiao, Chi-Yuan, et autres
Publié: (2026)
par: Hsiao, Chi-Yuan, et autres
Publié: (2026)
SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data
par: Wang, Hsuan-Fu, et autres
Publié: (2024)
par: Wang, Hsuan-Fu, et autres
Publié: (2024)
ML-SUPERB 2.0: Benchmarking Multilingual Speech Models Across Modeling Constraints, Languages, and Datasets
par: Shi, Jiatong, et autres
Publié: (2024)
par: Shi, Jiatong, et autres
Publié: (2024)
Parallel Synthesis for Autoregressive Speech Generation
par: Hsu, Po-chun, et autres
Publié: (2022)
par: Hsu, Po-chun, et autres
Publié: (2022)
SUTA-LM: Bridging Test-Time Adaptation and Language Model Rescoring for Robust ASR
par: Huang, Wei-Ping, et autres
Publié: (2025)
par: Huang, Wei-Ping, et autres
Publié: (2025)
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
par: Peng, Puyuan, et autres
Publié: (2025)
par: Peng, Puyuan, et autres
Publié: (2025)
VISinger2+: End-to-End Singing Voice Synthesis Augmented by Self-Supervised Learning Representation
par: Yu, Yifeng, et autres
Publié: (2024)
par: Yu, Yifeng, et autres
Publié: (2024)
Improving Speech Enhancement with Multi-Metric Supervision from Learned Quality Assessment
par: Wang, Wei, et autres
Publié: (2025)
par: Wang, Wei, et autres
Publié: (2025)
Enhancing Multilingual ASR for Unseen Languages via Language Embedding Modeling
par: Huang, Shao-Syuan, et autres
Publié: (2024)
par: Huang, Shao-Syuan, et autres
Publié: (2024)
Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction
par: Shi, Jiatong, et autres
Publié: (2023)
par: Shi, Jiatong, et autres
Publié: (2023)
Property Neurons in Self-Supervised Speech Transformers
par: Lin, Tzu-Quan, et autres
Publié: (2024)
par: Lin, Tzu-Quan, et autres
Publié: (2024)
PEFT for Speech: Unveiling Optimal Placement, Merging Strategies, and Ensemble Techniques
par: Lin, Tzu-Han, et autres
Publié: (2024)
par: Lin, Tzu-Han, et autres
Publié: (2024)
Leave No Knowledge Behind During Knowledge Distillation: Towards Practical and Effective Knowledge Distillation for Code-Switching ASR Using Realistic Data
par: Tseng, Liang-Hsuan, et autres
Publié: (2024)
par: Tseng, Liang-Hsuan, et autres
Publié: (2024)
Improving the Adversarial Robustness for Speaker Verification by Self-Supervised Learning
par: Wu, Haibin, et autres
Publié: (2021)
par: Wu, Haibin, et autres
Publié: (2021)
Probing the Robustness Properties of Neural Speech Codecs
par: Tseng, Wei-Cheng, et autres
Publié: (2025)
par: Tseng, Wei-Cheng, et autres
Publié: (2025)
From Alignment to Advancement: Bootstrapping Audio-Language Alignment with Synthetic Data
par: Kuan, Chun-Yi, et autres
Publié: (2025)
par: Kuan, Chun-Yi, et autres
Publié: (2025)
Dataset-Distillation Generative Model for Speech Emotion Recognition
par: Ritter-Gutierrez, Fabian, et autres
Publié: (2024)
par: Ritter-Gutierrez, Fabian, et autres
Publié: (2024)
Understanding Sounds, Missing the Questions: The Challenge of Object Hallucination in Large Audio-Language Models
par: Kuan, Chun-Yi, et autres
Publié: (2024)
par: Kuan, Chun-Yi, et autres
Publié: (2024)
TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
par: Tseng, Liang-Hsuan, et autres
Publié: (2025)
par: Tseng, Liang-Hsuan, et autres
Publié: (2025)
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
par: Yang, Chih-Kai, et autres
Publié: (2025)
par: Yang, Chih-Kai, et autres
Publié: (2025)
Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
par: Lin, Hsi-Che, et autres
Publié: (2024)
par: Lin, Hsi-Che, et autres
Publié: (2024)
Findings of the 2023 ML-SUPERB Challenge: Pre-Training and Evaluation over More Languages and Beyond
par: Shi, Jiatong, et autres
Publié: (2023)
par: Shi, Jiatong, et autres
Publié: (2023)
Speech-Copilot: Leveraging Large Language Models for Speech Processing via Task Decomposition, Modularization, and Program Generation
par: Kuan, Chun-Yi, et autres
Publié: (2024)
par: Kuan, Chun-Yi, et autres
Publié: (2024)
CodecFake: Enhancing Anti-Spoofing Models Against Deepfake Audios from Codec-Based Speech Synthesis Systems
par: Wu, Haibin, et autres
Publié: (2024)
par: Wu, Haibin, et autres
Publié: (2024)
SLAP: Learning Speaker and Health-Related Representations from Natural Language Supervision
par: Ando, Angelika, et autres
Publié: (2025)
par: Ando, Angelika, et autres
Publié: (2025)
CLaMP 3: Universal Music Information Retrieval Across Unaligned Modalities and Unseen Languages
par: Wu, Shangda, et autres
Publié: (2025)
par: Wu, Shangda, et autres
Publié: (2025)
Is Transfer Learning Necessary for Violin Transcription?
par: Peng, Yueh-Po, et autres
Publié: (2025)
par: Peng, Yueh-Po, et autres
Publié: (2025)
IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling
par: Huang, Kuan-Po, et autres
Publié: (2025)
par: Huang, Kuan-Po, et autres
Publié: (2025)
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation
par: Huang, Kuan-Po, et autres
Publié: (2026)
par: Huang, Kuan-Po, et autres
Publié: (2026)
Documents similaires
-
Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages
par: Huang, Kuan-Po, et autres
Publié: (2023) -
Low-Resource Self-Supervised Learning with SSL-Enhanced TTS
par: Hsu, Po-chun, et autres
Publié: (2023) -
Do Neural Codecs Generalize? A Controlled Study Across Unseen Languages and Non-Speech Tasks
par: Wang, Shih-Heng, et autres
Publié: (2026) -
Interface Design for Self-Supervised Speech Models
par: Shih, Yi-Jen, et autres
Publié: (2024) -
SMILE: Speech Meta In-Context Learning for Low-Resource Language Automatic Speech Recognition
par: Hsu, Ming-Hao, et autres
Publié: (2024)