Approaching Dialogue State Tracking via Aligning Speech Encoders and LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sedláček, Šimon, Yusuf, Bolaji, Švec, Ján, Hegde, Pradyoth, Kesiraju, Santosh, Plchot, Oldřich, Černocký, Jan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916787998687232
author Sedláček, Šimon
Yusuf, Bolaji
Švec, Ján
Hegde, Pradyoth
Kesiraju, Santosh
Plchot, Oldřich
Černocký, Jan
author_facet Sedláček, Šimon
Yusuf, Bolaji
Švec, Ján
Hegde, Pradyoth
Kesiraju, Santosh
Plchot, Oldřich
Černocký, Jan
contents In this work, we approach spoken Dialogue State Tracking (DST) by bridging the representation spaces of speech encoders and LLMs via a small connector module, with a focus on fully open-sourced and open-data components (WavLM-large, OLMo). We focus on ablating different aspects of such systems including full/LoRA adapter fine-tuning, the effect of agent turns in the dialogue history, as well as fuzzy matching-based output post-processing, which greatly improves performance of our systems on named entities in the dialogue slot values. We conduct our experiments on the SpokenWOZ dataset, and additionally utilize the Speech-Aware MultiWOZ dataset to augment our training data. Ultimately, our best-performing WavLM + connector + OLMo-1B aligned models achieve state of the art on the SpokenWOZ test set (34.66% JGA), and our system with Gemma-2-9B-instruct further surpasses this result, reaching 42.17% JGA on SpokenWOZ test.
format Preprint
id arxiv_https___arxiv_org_abs_2506_08633
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Approaching Dialogue State Tracking via Aligning Speech Encoders and LLMs
Sedláček, Šimon
Yusuf, Bolaji
Švec, Ján
Hegde, Pradyoth
Kesiraju, Santosh
Plchot, Oldřich
Černocký, Jan
Audio and Speech Processing
Computation and Language
In this work, we approach spoken Dialogue State Tracking (DST) by bridging the representation spaces of speech encoders and LLMs via a small connector module, with a focus on fully open-sourced and open-data components (WavLM-large, OLMo). We focus on ablating different aspects of such systems including full/LoRA adapter fine-tuning, the effect of agent turns in the dialogue history, as well as fuzzy matching-based output post-processing, which greatly improves performance of our systems on named entities in the dialogue slot values. We conduct our experiments on the SpokenWOZ dataset, and additionally utilize the Speech-Aware MultiWOZ dataset to augment our training data. Ultimately, our best-performing WavLM + connector + OLMo-1B aligned models achieve state of the art on the SpokenWOZ test set (34.66% JGA), and our system with Gemma-2-9B-instruct further surpasses this result, reaching 42.17% JGA on SpokenWOZ test.
title Approaching Dialogue State Tracking via Aligning Speech Encoders and LLMs
topic Audio and Speech Processing
Computation and Language
url https://arxiv.org/abs/2506.08633