AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Suglia, Alessandro, Greco, Claudio, Baker, Katie, Part, Jose L., Papaioannou, Ioannis, Eshghi, Arash, Konstas, Ioannis, Lemon, Oliver
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909228683231232
author Suglia, Alessandro
Greco, Claudio
Baker, Katie
Part, Jose L.
Papaioannou, Ioannis
Eshghi, Arash
Konstas, Ioannis
Lemon, Oliver
author_facet Suglia, Alessandro
Greco, Claudio
Baker, Katie
Part, Jose L.
Papaioannou, Ioannis
Eshghi, Arash
Konstas, Ioannis
Lemon, Oliver
contents AI personal assistants deployed via robots or wearables require embodied understanding to collaborate with humans effectively. However, current Vision-Language Models (VLMs) primarily focus on third-person view videos, neglecting the richness of egocentric perceptual experience. To address this gap, we propose three key contributions. First, we introduce the Egocentric Video Understanding Dataset (EVUD) for training VLMs on video captioning and question answering tasks specific to egocentric videos. Second, we present AlanaVLM, a 7B parameter VLM trained using parameter-efficient methods on EVUD. Finally, we evaluate AlanaVLM's capabilities on OpenEQA, a challenging benchmark for embodied video question answering. Our model achieves state-of-the-art performance, outperforming open-source models including strong Socratic models using GPT-4 as a planner by 3.6%. Additionally, we outperform Claude 3 and Gemini Pro Vision 1.0 and showcase competitive results compared to Gemini Pro 1.5 and GPT-4V, even surpassing the latter in spatial reasoning. This research paves the way for building efficient VLMs that can be deployed in robots or wearables, leveraging embodied video understanding to collaborate seamlessly with humans in everyday tasks, contributing to the next generation of Embodied AI.
format Preprint
id arxiv_https___arxiv_org_abs_2406_13807
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding
Suglia, Alessandro
Greco, Claudio
Baker, Katie
Part, Jose L.
Papaioannou, Ioannis
Eshghi, Arash
Konstas, Ioannis
Lemon, Oliver
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
AI personal assistants deployed via robots or wearables require embodied understanding to collaborate with humans effectively. However, current Vision-Language Models (VLMs) primarily focus on third-person view videos, neglecting the richness of egocentric perceptual experience. To address this gap, we propose three key contributions. First, we introduce the Egocentric Video Understanding Dataset (EVUD) for training VLMs on video captioning and question answering tasks specific to egocentric videos. Second, we present AlanaVLM, a 7B parameter VLM trained using parameter-efficient methods on EVUD. Finally, we evaluate AlanaVLM's capabilities on OpenEQA, a challenging benchmark for embodied video question answering. Our model achieves state-of-the-art performance, outperforming open-source models including strong Socratic models using GPT-4 as a planner by 3.6%. Additionally, we outperform Claude 3 and Gemini Pro Vision 1.0 and showcase competitive results compared to Gemini Pro 1.5 and GPT-4V, even surpassing the latter in spatial reasoning. This research paves the way for building efficient VLMs that can be deployed in robots or wearables, leveraging embodied video understanding to collaborate seamlessly with humans in everyday tasks, contributing to the next generation of Embodied AI.
title AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2406.13807