Device-Directed Speech Detection for Follow-up Conversations Using Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ognjen, Rudovic, Dighe, Pranay, Su, Yi, Garg, Vineet, Dharur, Sameer, Niu, Xiaochuan, Abdelaziz, Ahmed H., Adya, Saurabh, Tewfik, Ahmed
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913571650142208
author Ognjen
Rudovic
Dighe, Pranay
Su, Yi
Garg, Vineet
Dharur, Sameer
Niu, Xiaochuan
Abdelaziz, Ahmed H.
Adya, Saurabh
Tewfik, Ahmed
author_facet Ognjen
Rudovic
Dighe, Pranay
Su, Yi
Garg, Vineet
Dharur, Sameer
Niu, Xiaochuan
Abdelaziz, Ahmed H.
Adya, Saurabh
Tewfik, Ahmed
contents Follow-up conversations with virtual assistants (VAs) enable a user to seamlessly interact with a VA without the need to repeatedly invoke it using a keyword (after the first query). Therefore, accurate Device-directed Speech Detection (DDSD) from the follow-up queries is critical for enabling naturalistic user experience. To this end, we explore the notion of Large Language Models (LLMs) and model the first query when making inference about the follow-ups (based on the ASR-decoded text), via prompting of a pretrained LLM, or by adapting a binary classifier on top of the LLM. In doing so, we also exploit the ASR uncertainty when designing the LLM prompts. We show on the real-world dataset of follow-up conversations that this approach yields large gains (20-40% reduction in false alarms at 10% fixed false rejects) due to the joint modeling of the previous speech context and ASR uncertainty, compared to when follow-ups are modeled alone.
format Preprint
id arxiv_https___arxiv_org_abs_2411_00023
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Device-Directed Speech Detection for Follow-up Conversations Using Large Language Models
Ognjen
Rudovic
Dighe, Pranay
Su, Yi
Garg, Vineet
Dharur, Sameer
Niu, Xiaochuan
Abdelaziz, Ahmed H.
Adya, Saurabh
Tewfik, Ahmed
Audio and Speech Processing
Artificial Intelligence
Computation and Language
Sound
Follow-up conversations with virtual assistants (VAs) enable a user to seamlessly interact with a VA without the need to repeatedly invoke it using a keyword (after the first query). Therefore, accurate Device-directed Speech Detection (DDSD) from the follow-up queries is critical for enabling naturalistic user experience. To this end, we explore the notion of Large Language Models (LLMs) and model the first query when making inference about the follow-ups (based on the ASR-decoded text), via prompting of a pretrained LLM, or by adapting a binary classifier on top of the LLM. In doing so, we also exploit the ASR uncertainty when designing the LLM prompts. We show on the real-world dataset of follow-up conversations that this approach yields large gains (20-40% reduction in false alarms at 10% fixed false rejects) due to the joint modeling of the previous speech context and ASR uncertainty, compared to when follow-ups are modeled alone.
title Device-Directed Speech Detection for Follow-up Conversations Using Large Language Models
topic Audio and Speech Processing
Artificial Intelligence
Computation and Language
Sound
url https://arxiv.org/abs/2411.00023