User Interaction Patterns and Breakdowns in Conversing with LLM-Powered Voice Assistants

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mahmood, Amama, Wang, Junxiang, Yao, Bingsheng, Wang, Dakuo, Huang, Chien-Ming
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909406805884928
author Mahmood, Amama
Wang, Junxiang
Yao, Bingsheng
Wang, Dakuo
Huang, Chien-Ming
author_facet Mahmood, Amama
Wang, Junxiang
Yao, Bingsheng
Wang, Dakuo
Huang, Chien-Ming
contents Conventional Voice Assistants (VAs) rely on traditional language models to discern user intent and respond to their queries, leading to interactions that often lack a broader contextual understanding, an area in which Large Language Models (LLMs) excel. However, current LLMs are largely designed for text-based interactions, thus making it unclear how user interactions will evolve if their modality is changed to voice. In this work, we investigate whether LLMs can enrich VA interactions via an exploratory study with participants (N=20) using a ChatGPT-powered VA for three scenarios (medical self-diagnosis, creative planning, and discussion) with varied constraints, stakes, and objectivity. We observe that LLM-powered VA elicits richer interaction patterns that vary across tasks, showing its versatility. Notably, LLMs absorb the majority of VA intent recognition failures. We additionally discuss the potential of harnessing LLMs for more resilient and fluid user-VA interactions and provide design guidelines for tailoring LLMs for voice assistance.
format Preprint
id arxiv_https___arxiv_org_abs_2309_13879
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle User Interaction Patterns and Breakdowns in Conversing with LLM-Powered Voice Assistants
Mahmood, Amama
Wang, Junxiang
Yao, Bingsheng
Wang, Dakuo
Huang, Chien-Ming
Human-Computer Interaction
Conventional Voice Assistants (VAs) rely on traditional language models to discern user intent and respond to their queries, leading to interactions that often lack a broader contextual understanding, an area in which Large Language Models (LLMs) excel. However, current LLMs are largely designed for text-based interactions, thus making it unclear how user interactions will evolve if their modality is changed to voice. In this work, we investigate whether LLMs can enrich VA interactions via an exploratory study with participants (N=20) using a ChatGPT-powered VA for three scenarios (medical self-diagnosis, creative planning, and discussion) with varied constraints, stakes, and objectivity. We observe that LLM-powered VA elicits richer interaction patterns that vary across tasks, showing its versatility. Notably, LLMs absorb the majority of VA intent recognition failures. We additionally discuss the potential of harnessing LLMs for more resilient and fluid user-VA interactions and provide design guidelines for tailoring LLMs for voice assistance.
title User Interaction Patterns and Breakdowns in Conversing with LLM-Powered Voice Assistants
topic Human-Computer Interaction
url https://arxiv.org/abs/2309.13879