Exploring Fine-Tuning of Large Audio Language Models for Spoken Language Understanding under Limited Speech Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Choi, Youngwon, Jung, Jaeyoon, Kim, Hyeonyu, Nguyen, Huu-Kim, Kim, Hwayeon
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915741935075328
author Choi, Youngwon
Jung, Jaeyoon
Kim, Hyeonyu
Nguyen, Huu-Kim
Kim, Hwayeon
author_facet Choi, Youngwon
Jung, Jaeyoon
Kim, Hyeonyu
Nguyen, Huu-Kim
Kim, Hwayeon
contents Large Audio Language Models (LALMs) have emerged as powerful tools for speech-related tasks but remain underexplored for fine-tuning, especially with limited speech data. To bridge this gap, we systematically examine how different fine-tuning schemes including text-only, direct mixing, and curriculum learning affect spoken language understanding (SLU), focusing on scenarios where text-label pairs are abundant while paired speech-label data are limited. Results show that LALMs already achieve competitive performance with text-only fine-tuning, highlighting their strong generalization ability. Adding even small amounts of speech data (2-5%) yields substantial further gains, with curriculum learning particularly effective under scarce data. In cross-lingual SLU, combining source-language speech data with target-language text and minimal target-language speech data enables effective adaptation. Overall, this study provides practical insights into the LALM fine-tuning under realistic data constraints.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15389
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring Fine-Tuning of Large Audio Language Models for Spoken Language Understanding under Limited Speech Data
Choi, Youngwon
Jung, Jaeyoon
Kim, Hyeonyu
Nguyen, Huu-Kim
Kim, Hwayeon
Sound
Computation and Language
Machine Learning
Audio and Speech Processing
Large Audio Language Models (LALMs) have emerged as powerful tools for speech-related tasks but remain underexplored for fine-tuning, especially with limited speech data. To bridge this gap, we systematically examine how different fine-tuning schemes including text-only, direct mixing, and curriculum learning affect spoken language understanding (SLU), focusing on scenarios where text-label pairs are abundant while paired speech-label data are limited. Results show that LALMs already achieve competitive performance with text-only fine-tuning, highlighting their strong generalization ability. Adding even small amounts of speech data (2-5%) yields substantial further gains, with curriculum learning particularly effective under scarce data. In cross-lingual SLU, combining source-language speech data with target-language text and minimal target-language speech data enables effective adaptation. Overall, this study provides practical insights into the LALM fine-tuning under realistic data constraints.
title Exploring Fine-Tuning of Large Audio Language Models for Spoken Language Understanding under Limited Speech Data
topic Sound
Computation and Language
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2509.15389