I Was Blind but Now I See: Implementing Vision-Enabled Dialogue in Social Robots

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Abbo, Giulio Antonio, Belpaeme, Tony
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908312242487296
author Abbo, Giulio Antonio
Belpaeme, Tony
author_facet Abbo, Giulio Antonio
Belpaeme, Tony
contents In the rapidly evolving landscape of human-computer interaction, the integration of vision capabilities into conversational agents stands as a crucial advancement. This paper presents an initial implementation of a dialogue manager that leverages the latest progress in Large Language Models (e.g., GPT-4, IDEFICS) to enhance the traditional text-based prompts with real-time visual input. LLMs are used to interpret both textual prompts and visual stimuli, creating a more contextually aware conversational agent. The system's prompt engineering, incorporating dialogue with summarisation of the images, ensures a balance between context preservation and computational efficiency. Six interactions with a Furhat robot powered by this system are reported, illustrating and discussing the results obtained. By implementing this vision-enabled dialogue system, the paper envisions a future where conversational agents seamlessly blend textual and visual modalities, enabling richer, more context-aware dialogues.
format Preprint
id arxiv_https___arxiv_org_abs_2311_08957
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle I Was Blind but Now I See: Implementing Vision-Enabled Dialogue in Social Robots
Abbo, Giulio Antonio
Belpaeme, Tony
Robotics
Artificial Intelligence
Human-Computer Interaction
In the rapidly evolving landscape of human-computer interaction, the integration of vision capabilities into conversational agents stands as a crucial advancement. This paper presents an initial implementation of a dialogue manager that leverages the latest progress in Large Language Models (e.g., GPT-4, IDEFICS) to enhance the traditional text-based prompts with real-time visual input. LLMs are used to interpret both textual prompts and visual stimuli, creating a more contextually aware conversational agent. The system's prompt engineering, incorporating dialogue with summarisation of the images, ensures a balance between context preservation and computational efficiency. Six interactions with a Furhat robot powered by this system are reported, illustrating and discussing the results obtained. By implementing this vision-enabled dialogue system, the paper envisions a future where conversational agents seamlessly blend textual and visual modalities, enabling richer, more context-aware dialogues.
title I Was Blind but Now I See: Implementing Vision-Enabled Dialogue in Social Robots
topic Robotics
Artificial Intelligence
Human-Computer Interaction
url https://arxiv.org/abs/2311.08957