New Semantic Task for the French Spoken Language Understanding MEDIA Benchmark

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Alavoine, Nadège, Laperriere, Gaëlle, Servan, Christophe, Ghannay, Sahar, Rosset, Sophie
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911818762420224
author Alavoine, Nadège
Laperriere, Gaëlle
Servan, Christophe
Ghannay, Sahar
Rosset, Sophie
author_facet Alavoine, Nadège
Laperriere, Gaëlle
Servan, Christophe
Ghannay, Sahar
Rosset, Sophie
contents Intent classification and slot-filling are essential tasks of Spoken Language Understanding (SLU). In most SLUsystems, those tasks are realized by independent modules. For about fifteen years, models achieving both of themjointly and exploiting their mutual enhancement have been proposed. A multilingual module using a joint modelwas envisioned to create a touristic dialogue system for a European project, HumanE-AI-Net. A combination ofmultiple datasets, including the MEDIA dataset, was suggested for training this joint model. The MEDIA SLU datasetis a French dataset distributed since 2005 by ELRA, mainly used by the French research community and free foracademic research since 2020. Unfortunately, it is annotated only in slots but not intents. An enhanced version ofMEDIA annotated with intents has been built to extend its use to more tasks and use cases. This paper presents thesemi-automatic methodology used to obtain this enhanced version. In addition, we present the first results of SLUexperiments on this enhanced dataset using joint models for intent classification and slot-filling.
format Preprint
id arxiv_https___arxiv_org_abs_2403_19727
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle New Semantic Task for the French Spoken Language Understanding MEDIA Benchmark
Alavoine, Nadège
Laperriere, Gaëlle
Servan, Christophe
Ghannay, Sahar
Rosset, Sophie
Computation and Language
Artificial Intelligence
Intent classification and slot-filling are essential tasks of Spoken Language Understanding (SLU). In most SLUsystems, those tasks are realized by independent modules. For about fifteen years, models achieving both of themjointly and exploiting their mutual enhancement have been proposed. A multilingual module using a joint modelwas envisioned to create a touristic dialogue system for a European project, HumanE-AI-Net. A combination ofmultiple datasets, including the MEDIA dataset, was suggested for training this joint model. The MEDIA SLU datasetis a French dataset distributed since 2005 by ELRA, mainly used by the French research community and free foracademic research since 2020. Unfortunately, it is annotated only in slots but not intents. An enhanced version ofMEDIA annotated with intents has been built to extend its use to more tasks and use cases. This paper presents thesemi-automatic methodology used to obtain this enhanced version. In addition, we present the first results of SLUexperiments on this enhanced dataset using joint models for intent classification and slot-filling.
title New Semantic Task for the French Spoken Language Understanding MEDIA Benchmark
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2403.19727