I Was Blind but Now I See: Implementing Vision-Enabled Dialogue in Social Robots
Fuente:
arXiv
Saved in:
| Main Authors: | , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908312242487296 |
|---|---|
| author | Abbo, Giulio Antonio Belpaeme, Tony |
| author_facet | Abbo, Giulio Antonio Belpaeme, Tony |
| contents | In the rapidly evolving landscape of human-computer interaction, the integration of vision capabilities into conversational agents stands as a crucial advancement. This paper presents an initial implementation of a dialogue manager that leverages the latest progress in Large Language Models (e.g., GPT-4, IDEFICS) to enhance the traditional text-based prompts with real-time visual input. LLMs are used to interpret both textual prompts and visual stimuli, creating a more contextually aware conversational agent. The system's prompt engineering, incorporating dialogue with summarisation of the images, ensures a balance between context preservation and computational efficiency. Six interactions with a Furhat robot powered by this system are reported, illustrating and discussing the results obtained. By implementing this vision-enabled dialogue system, the paper envisions a future where conversational agents seamlessly blend textual and visual modalities, enabling richer, more context-aware dialogues. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2311_08957 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | I Was Blind but Now I See: Implementing Vision-Enabled Dialogue in Social Robots Abbo, Giulio Antonio Belpaeme, Tony Robotics Artificial Intelligence Human-Computer Interaction In the rapidly evolving landscape of human-computer interaction, the integration of vision capabilities into conversational agents stands as a crucial advancement. This paper presents an initial implementation of a dialogue manager that leverages the latest progress in Large Language Models (e.g., GPT-4, IDEFICS) to enhance the traditional text-based prompts with real-time visual input. LLMs are used to interpret both textual prompts and visual stimuli, creating a more contextually aware conversational agent. The system's prompt engineering, incorporating dialogue with summarisation of the images, ensures a balance between context preservation and computational efficiency. Six interactions with a Furhat robot powered by this system are reported, illustrating and discussing the results obtained. By implementing this vision-enabled dialogue system, the paper envisions a future where conversational agents seamlessly blend textual and visual modalities, enabling richer, more context-aware dialogues. |
| title | I Was Blind but Now I See: Implementing Vision-Enabled Dialogue in Social Robots |
| topic | Robotics Artificial Intelligence Human-Computer Interaction |
| url | https://arxiv.org/abs/2311.08957 |