Takeaways from Applying LLM Capabilities to Multiple Conversational Avatars in a VR Pilot Study

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Maslych, Mykola, Pumarada, Christian, Ghasemaghaei, Amirpouya, LaViola Jr, Joseph J.
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915092815151104
author Maslych, Mykola
Pumarada, Christian
Ghasemaghaei, Amirpouya
LaViola Jr, Joseph J.
author_facet Maslych, Mykola
Pumarada, Christian
Ghasemaghaei, Amirpouya
LaViola Jr, Joseph J.
contents We present a virtual reality (VR) environment featuring conversational avatars powered by a locally-deployed LLM, integrated with automatic speech recognition (ASR), text-to-speech (TTS), and lip-syncing. Through a pilot study, we explored the effects of three types of avatar status indicators during response generation. Our findings reveal design considerations for improving responsiveness and realism in LLM-driven conversational systems. We also detail two system architectures: one using an LLM-based state machine to control avatar behavior and another integrating retrieval-augmented generation (RAG) for context-grounded responses. Together, these contributions offer practical insights to guide future work in developing task-oriented conversational AI in VR environments.
format Preprint
id arxiv_https___arxiv_org_abs_2501_00168
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Takeaways from Applying LLM Capabilities to Multiple Conversational Avatars in a VR Pilot Study
Maslych, Mykola
Pumarada, Christian
Ghasemaghaei, Amirpouya
LaViola Jr, Joseph J.
Human-Computer Interaction
We present a virtual reality (VR) environment featuring conversational avatars powered by a locally-deployed LLM, integrated with automatic speech recognition (ASR), text-to-speech (TTS), and lip-syncing. Through a pilot study, we explored the effects of three types of avatar status indicators during response generation. Our findings reveal design considerations for improving responsiveness and realism in LLM-driven conversational systems. We also detail two system architectures: one using an LLM-based state machine to control avatar behavior and another integrating retrieval-augmented generation (RAG) for context-grounded responses. Together, these contributions offer practical insights to guide future work in developing task-oriented conversational AI in VR environments.
title Takeaways from Applying LLM Capabilities to Multiple Conversational Avatars in a VR Pilot Study
topic Human-Computer Interaction
url https://arxiv.org/abs/2501.00168