Speech-MASSIVE: A Multilingual Speech Dataset for SLU and Beyond

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Beomseok, Calapodescu, Ioan, Gaido, Marco, Negri, Matteo, Besacier, Laurent
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917743383543808
author Lee, Beomseok
Calapodescu, Ioan
Gaido, Marco
Negri, Matteo
Besacier, Laurent
author_facet Lee, Beomseok
Calapodescu, Ioan
Gaido, Marco
Negri, Matteo
Besacier, Laurent
contents We present Speech-MASSIVE, a multilingual Spoken Language Understanding (SLU) dataset comprising the speech counterpart for a portion of the MASSIVE textual corpus. Speech-MASSIVE covers 12 languages from different families and inherits from MASSIVE the annotations for the intent prediction and slot-filling tasks. Our extension is prompted by the scarcity of massively multilingual SLU datasets and the growing need for versatile speech datasets to assess foundation models (LLMs, speech encoders) across languages and tasks. We provide a multimodal, multitask, multilingual dataset and report SLU baselines using both cascaded and end-to-end architectures in various training scenarios (zero-shot, few-shot, and full fine-tune). Furthermore, we demonstrate the suitability of Speech-MASSIVE for benchmarking other tasks such as speech transcription, language identification, and speech translation. The dataset, models, and code are publicly available at: https://github.com/hlt-mt/Speech-MASSIVE
format Preprint
id arxiv_https___arxiv_org_abs_2408_03900
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Speech-MASSIVE: A Multilingual Speech Dataset for SLU and Beyond
Lee, Beomseok
Calapodescu, Ioan
Gaido, Marco
Negri, Matteo
Besacier, Laurent
Computation and Language
Sound
Audio and Speech Processing
We present Speech-MASSIVE, a multilingual Spoken Language Understanding (SLU) dataset comprising the speech counterpart for a portion of the MASSIVE textual corpus. Speech-MASSIVE covers 12 languages from different families and inherits from MASSIVE the annotations for the intent prediction and slot-filling tasks. Our extension is prompted by the scarcity of massively multilingual SLU datasets and the growing need for versatile speech datasets to assess foundation models (LLMs, speech encoders) across languages and tasks. We provide a multimodal, multitask, multilingual dataset and report SLU baselines using both cascaded and end-to-end architectures in various training scenarios (zero-shot, few-shot, and full fine-tune). Furthermore, we demonstrate the suitability of Speech-MASSIVE for benchmarking other tasks such as speech transcription, language identification, and speech translation. The dataset, models, and code are publicly available at: https://github.com/hlt-mt/Speech-MASSIVE
title Speech-MASSIVE: A Multilingual Speech Dataset for SLU and Beyond
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2408.03900