Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2403.14438 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917622519431168 |
|---|---|
| author | Wagner, Dominik Churchill, Alexander Sigtia, Siddharth Georgiou, Panayiotis Mirsamadi, Matt Mishra, Aarshee Marchi, Erik |
| author_facet | Wagner, Dominik Churchill, Alexander Sigtia, Siddharth Georgiou, Panayiotis Mirsamadi, Matt Mishra, Aarshee Marchi, Erik |
| contents | Interactions with virtual assistants typically start with a predefined trigger phrase followed by the user command. To make interactions with the assistant more intuitive, we explore whether it is feasible to drop the requirement that users must begin each command with a trigger phrase. We explore this task in three ways: First, we train classifiers using only acoustic information obtained from the audio waveform. Second, we take the decoder outputs of an automatic speech recognition (ASR) system, such as 1-best hypotheses, as input features to a large language model (LLM). Finally, we explore a multimodal system that combines acoustic and lexical features, as well as ASR decoder signals in an LLM. Using multimodal information yields relative equal-error-rate improvements over text-only and audio-only models of up to 39% and 61%. Increasing the size of the LLM and training with low-rank adaption leads to further relative EER reductions of up to 18% on our dataset. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2403_14438 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | A Multimodal Approach to Device-Directed Speech Detection with Large Language Models Wagner, Dominik Churchill, Alexander Sigtia, Siddharth Georgiou, Panayiotis Mirsamadi, Matt Mishra, Aarshee Marchi, Erik Computation and Language Machine Learning Audio and Speech Processing Interactions with virtual assistants typically start with a predefined trigger phrase followed by the user command. To make interactions with the assistant more intuitive, we explore whether it is feasible to drop the requirement that users must begin each command with a trigger phrase. We explore this task in three ways: First, we train classifiers using only acoustic information obtained from the audio waveform. Second, we take the decoder outputs of an automatic speech recognition (ASR) system, such as 1-best hypotheses, as input features to a large language model (LLM). Finally, we explore a multimodal system that combines acoustic and lexical features, as well as ASR decoder signals in an LLM. Using multimodal information yields relative equal-error-rate improvements over text-only and audio-only models of up to 39% and 61%. Increasing the size of the LLM and training with low-rank adaption leads to further relative EER reductions of up to 18% on our dataset. |
| title | A Multimodal Approach to Device-Directed Speech Detection with Large Language Models |
| topic | Computation and Language Machine Learning Audio and Speech Processing |
| url | https://arxiv.org/abs/2403.14438 |