UniverSLU: Universal Spoken Language Understanding for Diverse Tasks with Natural Language Instructions

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Arora, Siddhant, Futami, Hayato, Jung, Jee-weon, Peng, Yifan, Sharma, Roshan, Kashiwagi, Yosuke, Tsunoo, Emiru, Livescu, Karen, Watanabe, Shinji
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910396916432896
author Arora, Siddhant
Futami, Hayato
Jung, Jee-weon
Peng, Yifan
Sharma, Roshan
Kashiwagi, Yosuke
Tsunoo, Emiru
Livescu, Karen
Watanabe, Shinji
author_facet Arora, Siddhant
Futami, Hayato
Jung, Jee-weon
Peng, Yifan
Sharma, Roshan
Kashiwagi, Yosuke
Tsunoo, Emiru
Livescu, Karen
Watanabe, Shinji
contents Recent studies leverage large language models with multi-tasking capabilities, using natural language prompts to guide the model's behavior and surpassing performance of task-specific models. Motivated by this, we ask: can we build a single model that jointly performs various spoken language understanding (SLU) tasks? We start by adapting a pre-trained automatic speech recognition model to additional tasks using single-token task specifiers. We enhance this approach through instruction tuning, i.e., finetuning by describing the task using natural language instructions followed by the list of label options. Our approach can generalize to new task descriptions for the seen tasks during inference, thereby enhancing its user-friendliness. We demonstrate the efficacy of our single multi-task learning model "UniverSLU" for 12 speech classification and sequence generation task types spanning 17 datasets and 9 languages. On most tasks, UniverSLU achieves competitive performance and often even surpasses task-specific models. Additionally, we assess the zero-shot capabilities, finding that the model generalizes to new datasets and languages for seen task types.
format Preprint
id arxiv_https___arxiv_org_abs_2310_02973
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle UniverSLU: Universal Spoken Language Understanding for Diverse Tasks with Natural Language Instructions
Arora, Siddhant
Futami, Hayato
Jung, Jee-weon
Peng, Yifan
Sharma, Roshan
Kashiwagi, Yosuke
Tsunoo, Emiru
Livescu, Karen
Watanabe, Shinji
Computation and Language
Sound
Audio and Speech Processing
Recent studies leverage large language models with multi-tasking capabilities, using natural language prompts to guide the model's behavior and surpassing performance of task-specific models. Motivated by this, we ask: can we build a single model that jointly performs various spoken language understanding (SLU) tasks? We start by adapting a pre-trained automatic speech recognition model to additional tasks using single-token task specifiers. We enhance this approach through instruction tuning, i.e., finetuning by describing the task using natural language instructions followed by the list of label options. Our approach can generalize to new task descriptions for the seen tasks during inference, thereby enhancing its user-friendliness. We demonstrate the efficacy of our single multi-task learning model "UniverSLU" for 12 speech classification and sequence generation task types spanning 17 datasets and 9 languages. On most tasks, UniverSLU achieves competitive performance and often even surpasses task-specific models. Additionally, we assess the zero-shot capabilities, finding that the model generalizes to new datasets and languages for seen task types.
title UniverSLU: Universal Spoken Language Understanding for Diverse Tasks with Natural Language Instructions
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2310.02973