SpeechLLMs for Large-scale Contextualized Zero-shot Slot Filling

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hacioglu, Kadri, E, Manjunath K, Stolcke, Andreas
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908599773560832
author Hacioglu, Kadri
E, Manjunath K
Stolcke, Andreas
author_facet Hacioglu, Kadri
E, Manjunath K
Stolcke, Andreas
contents Slot filling is a crucial subtask in spoken language understanding (SLU), traditionally implemented as a cascade of speech recognition followed by one or more natural language understanding (NLU) components. The recent advent of speech-based large language models (speechLLMs), which integrate speech and textual foundation models, has opened new avenues for achieving speech understanding tasks in a more unified, generative, and instruction-following manner while promising data and compute efficiency with zero-shot abilities, generalizing to unseen slot labels. We address the slot-filling task by creating an empirical upper bound for the task, identifying performance, robustness, and generalization gaps, and proposing improvements to the training data, architecture, and training strategies to narrow the gap with the upper bound result. We show that each of these measures improve performance substantially, while highlighting practical challenges and providing empirical guidance and insights for harnessing these emerging models.
format Preprint
id arxiv_https___arxiv_org_abs_2510_15851
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SpeechLLMs for Large-scale Contextualized Zero-shot Slot Filling
Hacioglu, Kadri
E, Manjunath K
Stolcke, Andreas
Computation and Language
Machine Learning
Slot filling is a crucial subtask in spoken language understanding (SLU), traditionally implemented as a cascade of speech recognition followed by one or more natural language understanding (NLU) components. The recent advent of speech-based large language models (speechLLMs), which integrate speech and textual foundation models, has opened new avenues for achieving speech understanding tasks in a more unified, generative, and instruction-following manner while promising data and compute efficiency with zero-shot abilities, generalizing to unseen slot labels. We address the slot-filling task by creating an empirical upper bound for the task, identifying performance, robustness, and generalization gaps, and proposing improvements to the training data, architecture, and training strategies to narrow the gap with the upper bound result. We show that each of these measures improve performance substantially, while highlighting practical challenges and providing empirical guidance and insights for harnessing these emerging models.
title SpeechLLMs for Large-scale Contextualized Zero-shot Slot Filling
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2510.15851