MAC-SLU: Multi-Intent Automotive Cabin Spoken Language Understanding Benchmark

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Peng, Yuezhang, Cai, Chonghao, Liu, Ziang, Fan, Shuai, Jiang, Sheng, Xu, Hua, Liu, Yuxin, Chen, Qiguang, Xu, Kele, Li, Yao, Wang, Sheng, Qin, Libo, Chen, Xie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917116303638528
author Peng, Yuezhang
Cai, Chonghao
Liu, Ziang
Fan, Shuai
Jiang, Sheng
Xu, Hua
Liu, Yuxin
Chen, Qiguang
Xu, Kele
Li, Yao
Wang, Sheng
Qin, Libo
Chen, Xie
author_facet Peng, Yuezhang
Cai, Chonghao
Liu, Ziang
Fan, Shuai
Jiang, Sheng
Xu, Hua
Liu, Yuxin
Chen, Qiguang
Xu, Kele
Li, Yao
Wang, Sheng
Qin, Libo
Chen, Xie
contents Spoken Language Understanding (SLU), which aims to extract user semantics to execute downstream tasks, is a crucial component of task-oriented dialog systems. Existing SLU datasets generally lack sufficient diversity and complexity, and there is an absence of a unified benchmark for the latest Large Language Models (LLMs) and Large Audio Language Models (LALMs). This work introduces MAC-SLU, a novel Multi-Intent Automotive Cabin Spoken Language Understanding Dataset, which increases the difficulty of the SLU task by incorporating authentic and complex multi-intent data. Based on MAC-SLU, we conducted a comprehensive benchmark of leading open-source LLMs and LALMs, covering methods like in-context learning, supervised fine-tuning (SFT), and end-to-end (E2E) and pipeline paradigms. Our experiments show that while LLMs and LALMs have the potential to complete SLU tasks through in-context learning, their performance still lags significantly behind SFT. Meanwhile, E2E LALMs demonstrate performance comparable to pipeline approaches and effectively avoid error propagation from speech recognition. Code\footnote{https://github.com/Gatsby-web/MAC\_SLU} and datasets\footnote{huggingface.co/datasets/Gatsby1984/MAC\_SLU} are released publicly.
format Preprint
id arxiv_https___arxiv_org_abs_2512_01603
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MAC-SLU: Multi-Intent Automotive Cabin Spoken Language Understanding Benchmark
Peng, Yuezhang
Cai, Chonghao
Liu, Ziang
Fan, Shuai
Jiang, Sheng
Xu, Hua
Liu, Yuxin
Chen, Qiguang
Xu, Kele
Li, Yao
Wang, Sheng
Qin, Libo
Chen, Xie
Computation and Language
Multimedia
Spoken Language Understanding (SLU), which aims to extract user semantics to execute downstream tasks, is a crucial component of task-oriented dialog systems. Existing SLU datasets generally lack sufficient diversity and complexity, and there is an absence of a unified benchmark for the latest Large Language Models (LLMs) and Large Audio Language Models (LALMs). This work introduces MAC-SLU, a novel Multi-Intent Automotive Cabin Spoken Language Understanding Dataset, which increases the difficulty of the SLU task by incorporating authentic and complex multi-intent data. Based on MAC-SLU, we conducted a comprehensive benchmark of leading open-source LLMs and LALMs, covering methods like in-context learning, supervised fine-tuning (SFT), and end-to-end (E2E) and pipeline paradigms. Our experiments show that while LLMs and LALMs have the potential to complete SLU tasks through in-context learning, their performance still lags significantly behind SFT. Meanwhile, E2E LALMs demonstrate performance comparable to pipeline approaches and effectively avoid error propagation from speech recognition. Code\footnote{https://github.com/Gatsby-web/MAC\_SLU} and datasets\footnote{huggingface.co/datasets/Gatsby1984/MAC\_SLU} are released publicly.
title MAC-SLU: Multi-Intent Automotive Cabin Spoken Language Understanding Benchmark
topic Computation and Language
Multimedia
url https://arxiv.org/abs/2512.01603