ITALIC: An Italian Intent Classification Dataset

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Koudounas, Alkis, La Quatra, Moreno, Vaiani, Lorenzo, Colomba, Luca, Attanasio, Giuseppe, Pastor, Eliana, Cagliero, Luca, Baralis, Elena
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911930098122752
author Koudounas, Alkis
La Quatra, Moreno
Vaiani, Lorenzo
Colomba, Luca
Attanasio, Giuseppe
Pastor, Eliana
Cagliero, Luca
Baralis, Elena
author_facet Koudounas, Alkis
La Quatra, Moreno
Vaiani, Lorenzo
Colomba, Luca
Attanasio, Giuseppe
Pastor, Eliana
Cagliero, Luca
Baralis, Elena
contents Recent large-scale Spoken Language Understanding datasets focus predominantly on English and do not account for language-specific phenomena such as particular phonemes or words in different lects. We introduce ITALIC, the first large-scale speech dataset designed for intent classification in Italian. The dataset comprises 16,521 crowdsourced audio samples recorded by 70 speakers from various Italian regions and annotated with intent labels and additional metadata. We explore the versatility of ITALIC by evaluating current state-of-the-art speech and text models. Results on intent classification suggest that increasing scale and running language adaptation yield better speech models, monolingual text models outscore multilingual ones, and that speech recognition on ITALIC is more challenging than on existing Italian benchmarks. We release both the dataset and the annotation scheme to streamline the development of new Italian SLU models and language-specific datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2306_08502
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle ITALIC: An Italian Intent Classification Dataset
Koudounas, Alkis
La Quatra, Moreno
Vaiani, Lorenzo
Colomba, Luca
Attanasio, Giuseppe
Pastor, Eliana
Cagliero, Luca
Baralis, Elena
Computation and Language
Sound
Audio and Speech Processing
Recent large-scale Spoken Language Understanding datasets focus predominantly on English and do not account for language-specific phenomena such as particular phonemes or words in different lects. We introduce ITALIC, the first large-scale speech dataset designed for intent classification in Italian. The dataset comprises 16,521 crowdsourced audio samples recorded by 70 speakers from various Italian regions and annotated with intent labels and additional metadata. We explore the versatility of ITALIC by evaluating current state-of-the-art speech and text models. Results on intent classification suggest that increasing scale and running language adaptation yield better speech models, monolingual text models outscore multilingual ones, and that speech recognition on ITALIC is more challenging than on existing Italian benchmarks. We release both the dataset and the annotation scheme to streamline the development of new Italian SLU models and language-specific datasets.
title ITALIC: An Italian Intent Classification Dataset
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2306.08502