MAC-SLU: Multi-Intent Automotive Cabin Spoken Language Understanding Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917116303638528 |
|---|---|
| author | Peng, Yuezhang Cai, Chonghao Liu, Ziang Fan, Shuai Jiang, Sheng Xu, Hua Liu, Yuxin Chen, Qiguang Xu, Kele Li, Yao Wang, Sheng Qin, Libo Chen, Xie |
| author_facet | Peng, Yuezhang Cai, Chonghao Liu, Ziang Fan, Shuai Jiang, Sheng Xu, Hua Liu, Yuxin Chen, Qiguang Xu, Kele Li, Yao Wang, Sheng Qin, Libo Chen, Xie |
| contents | Spoken Language Understanding (SLU), which aims to extract user semantics to execute downstream tasks, is a crucial component of task-oriented dialog systems. Existing SLU datasets generally lack sufficient diversity and complexity, and there is an absence of a unified benchmark for the latest Large Language Models (LLMs) and Large Audio Language Models (LALMs). This work introduces MAC-SLU, a novel Multi-Intent Automotive Cabin Spoken Language Understanding Dataset, which increases the difficulty of the SLU task by incorporating authentic and complex multi-intent data. Based on MAC-SLU, we conducted a comprehensive benchmark of leading open-source LLMs and LALMs, covering methods like in-context learning, supervised fine-tuning (SFT), and end-to-end (E2E) and pipeline paradigms. Our experiments show that while LLMs and LALMs have the potential to complete SLU tasks through in-context learning, their performance still lags significantly behind SFT. Meanwhile, E2E LALMs demonstrate performance comparable to pipeline approaches and effectively avoid error propagation from speech recognition. Code\footnote{https://github.com/Gatsby-web/MAC\_SLU} and datasets\footnote{huggingface.co/datasets/Gatsby1984/MAC\_SLU} are released publicly. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_01603 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MAC-SLU: Multi-Intent Automotive Cabin Spoken Language Understanding Benchmark Peng, Yuezhang Cai, Chonghao Liu, Ziang Fan, Shuai Jiang, Sheng Xu, Hua Liu, Yuxin Chen, Qiguang Xu, Kele Li, Yao Wang, Sheng Qin, Libo Chen, Xie Computation and Language Multimedia Spoken Language Understanding (SLU), which aims to extract user semantics to execute downstream tasks, is a crucial component of task-oriented dialog systems. Existing SLU datasets generally lack sufficient diversity and complexity, and there is an absence of a unified benchmark for the latest Large Language Models (LLMs) and Large Audio Language Models (LALMs). This work introduces MAC-SLU, a novel Multi-Intent Automotive Cabin Spoken Language Understanding Dataset, which increases the difficulty of the SLU task by incorporating authentic and complex multi-intent data. Based on MAC-SLU, we conducted a comprehensive benchmark of leading open-source LLMs and LALMs, covering methods like in-context learning, supervised fine-tuning (SFT), and end-to-end (E2E) and pipeline paradigms. Our experiments show that while LLMs and LALMs have the potential to complete SLU tasks through in-context learning, their performance still lags significantly behind SFT. Meanwhile, E2E LALMs demonstrate performance comparable to pipeline approaches and effectively avoid error propagation from speech recognition. Code\footnote{https://github.com/Gatsby-web/MAC\_SLU} and datasets\footnote{huggingface.co/datasets/Gatsby1984/MAC\_SLU} are released publicly. |
| title | MAC-SLU: Multi-Intent Automotive Cabin Spoken Language Understanding Benchmark |
| topic | Computation and Language Multimedia |
| url | https://arxiv.org/abs/2512.01603 |