A Backpack Full of Skills: Egocentric Video Understanding with Diverse Task Perspectives

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Peirone, Simone Alberto, Pistilli, Francesca, Alliegro, Antonio, Averta, Giuseppe
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914703338373120
author Peirone, Simone Alberto
Pistilli, Francesca
Alliegro, Antonio
Averta, Giuseppe
author_facet Peirone, Simone Alberto
Pistilli, Francesca
Alliegro, Antonio
Averta, Giuseppe
contents Human comprehension of a video stream is naturally broad: in a few instants, we are able to understand what is happening, the relevance and relationship of objects, and forecast what will follow in the near future, everything all at once. We believe that - to effectively transfer such an holistic perception to intelligent machines - an important role is played by learning to correlate concepts and to abstract knowledge coming from different tasks, to synergistically exploit them when learning novel skills. To accomplish this, we seek for a unified approach to video understanding which combines shared temporal modelling of human actions with minimal overhead, to support multiple downstream tasks and enable cooperation when learning novel skills. We then propose EgoPack, a solution that creates a collection of task perspectives that can be carried across downstream tasks and used as a potential source of additional insights, as a backpack of skills that a robot can carry around and use when needed. We demonstrate the effectiveness and efficiency of our approach on four Ego4D benchmarks, outperforming current state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2403_03037
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Backpack Full of Skills: Egocentric Video Understanding with Diverse Task Perspectives
Peirone, Simone Alberto
Pistilli, Francesca
Alliegro, Antonio
Averta, Giuseppe
Computer Vision and Pattern Recognition
Machine Learning
Human comprehension of a video stream is naturally broad: in a few instants, we are able to understand what is happening, the relevance and relationship of objects, and forecast what will follow in the near future, everything all at once. We believe that - to effectively transfer such an holistic perception to intelligent machines - an important role is played by learning to correlate concepts and to abstract knowledge coming from different tasks, to synergistically exploit them when learning novel skills. To accomplish this, we seek for a unified approach to video understanding which combines shared temporal modelling of human actions with minimal overhead, to support multiple downstream tasks and enable cooperation when learning novel skills. We then propose EgoPack, a solution that creates a collection of task perspectives that can be carried across downstream tasks and used as a potential source of additional insights, as a backpack of skills that a robot can carry around and use when needed. We demonstrate the effectiveness and efficiency of our approach on four Ego4D benchmarks, outperforming current state-of-the-art methods.
title A Backpack Full of Skills: Egocentric Video Understanding with Diverse Task Perspectives
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2403.03037