SELMA: A Speech-Enabled Language Model for Virtual Assistant Interactions

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wagner, Dominik, Churchill, Alexander, Sigtia, Siddharth, Marchi, Erik
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917909700280320
author Wagner, Dominik
Churchill, Alexander
Sigtia, Siddharth
Marchi, Erik
author_facet Wagner, Dominik
Churchill, Alexander
Sigtia, Siddharth
Marchi, Erik
contents In this work, we present and evaluate SELMA, a Speech-Enabled Language Model for virtual Assistant interactions that integrates audio and text as inputs to a Large Language Model (LLM). SELMA is designed to handle three primary and two auxiliary tasks related to interactions with virtual assistants simultaneously within a single end-to-end model. We employ low-rank adaptation modules for parameter-efficient training of both the audio encoder and the LLM. Additionally, we implement a feature pooling strategy enabling the system to recognize global patterns and improve accuracy on tasks less reliant on individual sequence elements. Experimental results on Voice Trigger (VT) detection, Device-Directed Speech Detection (DDSD), and Automatic Speech Recognition (ASR), demonstrate that our approach both simplifies the typical input processing pipeline of virtual assistants significantly and also improves performance compared to dedicated models for each individual task. SELMA yields relative Equal-Error Rate improvements of 64% on the VT detection task, and 22% on DDSD, while also achieving word error rates close to the baseline.
format Preprint
id arxiv_https___arxiv_org_abs_2501_19377
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SELMA: A Speech-Enabled Language Model for Virtual Assistant Interactions
Wagner, Dominik
Churchill, Alexander
Sigtia, Siddharth
Marchi, Erik
Sound
Computation and Language
Machine Learning
Audio and Speech Processing
In this work, we present and evaluate SELMA, a Speech-Enabled Language Model for virtual Assistant interactions that integrates audio and text as inputs to a Large Language Model (LLM). SELMA is designed to handle three primary and two auxiliary tasks related to interactions with virtual assistants simultaneously within a single end-to-end model. We employ low-rank adaptation modules for parameter-efficient training of both the audio encoder and the LLM. Additionally, we implement a feature pooling strategy enabling the system to recognize global patterns and improve accuracy on tasks less reliant on individual sequence elements. Experimental results on Voice Trigger (VT) detection, Device-Directed Speech Detection (DDSD), and Automatic Speech Recognition (ASR), demonstrate that our approach both simplifies the typical input processing pipeline of virtual assistants significantly and also improves performance compared to dedicated models for each individual task. SELMA yields relative Equal-Error Rate improvements of 64% on the VT detection task, and 22% on DDSD, while also achieving word error rates close to the baseline.
title SELMA: A Speech-Enabled Language Model for Virtual Assistant Interactions
topic Sound
Computation and Language
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2501.19377