'What did the Robot do in my Absence?' Video Foundation Models to Enhance Intermittent Supervision

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Katuwandeniya, Kavindie, Tian, Leimin, Kulić, Dana
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910699584749568
author Katuwandeniya, Kavindie
Tian, Leimin
Kulić, Dana
author_facet Katuwandeniya, Kavindie
Tian, Leimin
Kulić, Dana
contents This paper investigates the application of Video Foundation Models (ViFMs) for generating robot data summaries to enhance intermittent human supervision of robot teams. We propose a novel framework that produces both generic and query-driven summaries of long-duration robot vision data in three modalities: storyboards, short videos, and text. Through a user study involving 30 participants, we evaluate the efficacy of these summary methods in allowing operators to accurately retrieve the observations and actions that occurred while the robot was operating without supervision over an extended duration (40 min). Our findings reveal that query-driven summaries significantly improve retrieval accuracy compared to generic summaries or raw data, albeit with increased task duration. Storyboards are found to be the most effective presentation modality, especially for object-related queries. This work represents, to our knowledge, the first zero-shot application of ViFMs for generating multi-modal robot-to-human communication in intermittent supervision contexts, demonstrating both the promise and limitations of these models in human-robot interaction (HRI) scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2411_10016
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle 'What did the Robot do in my Absence?' Video Foundation Models to Enhance Intermittent Supervision
Katuwandeniya, Kavindie
Tian, Leimin
Kulić, Dana
Robotics
This paper investigates the application of Video Foundation Models (ViFMs) for generating robot data summaries to enhance intermittent human supervision of robot teams. We propose a novel framework that produces both generic and query-driven summaries of long-duration robot vision data in three modalities: storyboards, short videos, and text. Through a user study involving 30 participants, we evaluate the efficacy of these summary methods in allowing operators to accurately retrieve the observations and actions that occurred while the robot was operating without supervision over an extended duration (40 min). Our findings reveal that query-driven summaries significantly improve retrieval accuracy compared to generic summaries or raw data, albeit with increased task duration. Storyboards are found to be the most effective presentation modality, especially for object-related queries. This work represents, to our knowledge, the first zero-shot application of ViFMs for generating multi-modal robot-to-human communication in intermittent supervision contexts, demonstrating both the promise and limitations of these models in human-robot interaction (HRI) scenarios.
title 'What did the Robot do in my Absence?' Video Foundation Models to Enhance Intermittent Supervision
topic Robotics
url https://arxiv.org/abs/2411.10016