Exploring Fine-Tuning of Large Audio Language Models for Spoken Language Understanding under Limited Speech Data
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915741935075328 |
|---|---|
| author | Choi, Youngwon Jung, Jaeyoon Kim, Hyeonyu Nguyen, Huu-Kim Kim, Hwayeon |
| author_facet | Choi, Youngwon Jung, Jaeyoon Kim, Hyeonyu Nguyen, Huu-Kim Kim, Hwayeon |
| contents | Large Audio Language Models (LALMs) have emerged as powerful tools for speech-related tasks but remain underexplored for fine-tuning, especially with limited speech data. To bridge this gap, we systematically examine how different fine-tuning schemes including text-only, direct mixing, and curriculum learning affect spoken language understanding (SLU), focusing on scenarios where text-label pairs are abundant while paired speech-label data are limited. Results show that LALMs already achieve competitive performance with text-only fine-tuning, highlighting their strong generalization ability. Adding even small amounts of speech data (2-5%) yields substantial further gains, with curriculum learning particularly effective under scarce data. In cross-lingual SLU, combining source-language speech data with target-language text and minimal target-language speech data enables effective adaptation. Overall, this study provides practical insights into the LALM fine-tuning under realistic data constraints. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_15389 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Exploring Fine-Tuning of Large Audio Language Models for Spoken Language Understanding under Limited Speech Data Choi, Youngwon Jung, Jaeyoon Kim, Hyeonyu Nguyen, Huu-Kim Kim, Hwayeon Sound Computation and Language Machine Learning Audio and Speech Processing Large Audio Language Models (LALMs) have emerged as powerful tools for speech-related tasks but remain underexplored for fine-tuning, especially with limited speech data. To bridge this gap, we systematically examine how different fine-tuning schemes including text-only, direct mixing, and curriculum learning affect spoken language understanding (SLU), focusing on scenarios where text-label pairs are abundant while paired speech-label data are limited. Results show that LALMs already achieve competitive performance with text-only fine-tuning, highlighting their strong generalization ability. Adding even small amounts of speech data (2-5%) yields substantial further gains, with curriculum learning particularly effective under scarce data. In cross-lingual SLU, combining source-language speech data with target-language text and minimal target-language speech data enables effective adaptation. Overall, this study provides practical insights into the LALM fine-tuning under realistic data constraints. |
| title | Exploring Fine-Tuning of Large Audio Language Models for Spoken Language Understanding under Limited Speech Data |
| topic | Sound Computation and Language Machine Learning Audio and Speech Processing |
| url | https://arxiv.org/abs/2509.15389 |