Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vendrame, Katia, Yusuf, Bolaji, Kesiraju, Santosh, Sedláček, Šimon, Plchot, Oldřich, Černocký, Jan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908679395082240
author Vendrame, Katia
Yusuf, Bolaji
Kesiraju, Santosh
Sedláček, Šimon
Plchot, Oldřich
Černocký, Jan
author_facet Vendrame, Katia
Yusuf, Bolaji
Kesiraju, Santosh
Sedláček, Šimon
Plchot, Oldřich
Černocký, Jan
contents End-to-end spoken dialogue state tracking (DST) is made difficult by the tandem of having to handle speech input and data scarcity. Combining speech foundation encoders and large language models has been proposed in recent work as to alleviate some of this difficulty. Although this approach has been shown to result in strong spoken DST models, achieving state-of-the-art performance in realistic multi-turn DST, it struggles to generalize across domains and requires annotated spoken DST training data for each domain of interest. However, collecting such data for every target domain is both costly and difficult. Noting that textual DST data is more easily obtained for various domains, in this work, we propose jointly training on available spoken DST data and written textual data from other domains as a way to achieve cross-domain generalization. We conduct experiments which show the efficacy of our proposed method for getting good cross-domain DST performance without relying on spoken training data from the target domains.
format Preprint
id arxiv_https___arxiv_org_abs_2511_22503
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
Vendrame, Katia
Yusuf, Bolaji
Kesiraju, Santosh
Sedláček, Šimon
Plchot, Oldřich
Černocký, Jan
Computation and Language
Sound
Audio and Speech Processing
End-to-end spoken dialogue state tracking (DST) is made difficult by the tandem of having to handle speech input and data scarcity. Combining speech foundation encoders and large language models has been proposed in recent work as to alleviate some of this difficulty. Although this approach has been shown to result in strong spoken DST models, achieving state-of-the-art performance in realistic multi-turn DST, it struggles to generalize across domains and requires annotated spoken DST training data for each domain of interest. However, collecting such data for every target domain is both costly and difficult. Noting that textual DST data is more easily obtained for various domains, in this work, we propose jointly training on available spoken DST data and written textual data from other domains as a way to achieve cross-domain generalization. We conduct experiments which show the efficacy of our proposed method for getting good cross-domain DST performance without relying on spoken training data from the target domains.
title Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2511.22503