Gated Temporal Diffusion for Stochastic Long-Term Dense Anticipation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zatsarynna, Olga, Bahrami, Emad, Farha, Yazan Abu, Francesca, Gianpiero, Gall, Juergen
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909258113613824
author Zatsarynna, Olga
Bahrami, Emad
Farha, Yazan Abu
Francesca, Gianpiero
Gall, Juergen
author_facet Zatsarynna, Olga
Bahrami, Emad
Farha, Yazan Abu
Francesca, Gianpiero
Gall, Juergen
contents Long-term action anticipation has become an important task for many applications such as autonomous driving and human-robot interaction. Unlike short-term anticipation, predicting more actions into the future imposes a real challenge with the increasing uncertainty in longer horizons. While there has been a significant progress in predicting more actions into the future, most of the proposed methods address the task in a deterministic setup and ignore the underlying uncertainty. In this paper, we propose a novel Gated Temporal Diffusion (GTD) network that models the uncertainty of both the observation and the future predictions. As generator, we introduce a Gated Anticipation Network (GTAN) to model both observed and unobserved frames of a video in a mutual representation. On the one hand, using a mutual representation for past and future allows us to jointly model ambiguities in the observation and future, while on the other hand GTAN can by design treat the observed and unobserved parts differently and steer the information flow between them. Our model achieves state-of-the-art results on the Breakfast, Assembly101 and 50Salads datasets in both stochastic and deterministic settings. Code: https://github.com/olga-zats/GTDA .
format Preprint
id arxiv_https___arxiv_org_abs_2407_11954
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Gated Temporal Diffusion for Stochastic Long-Term Dense Anticipation
Zatsarynna, Olga
Bahrami, Emad
Farha, Yazan Abu
Francesca, Gianpiero
Gall, Juergen
Computer Vision and Pattern Recognition
Long-term action anticipation has become an important task for many applications such as autonomous driving and human-robot interaction. Unlike short-term anticipation, predicting more actions into the future imposes a real challenge with the increasing uncertainty in longer horizons. While there has been a significant progress in predicting more actions into the future, most of the proposed methods address the task in a deterministic setup and ignore the underlying uncertainty. In this paper, we propose a novel Gated Temporal Diffusion (GTD) network that models the uncertainty of both the observation and the future predictions. As generator, we introduce a Gated Anticipation Network (GTAN) to model both observed and unobserved frames of a video in a mutual representation. On the one hand, using a mutual representation for past and future allows us to jointly model ambiguities in the observation and future, while on the other hand GTAN can by design treat the observed and unobserved parts differently and steer the information flow between them. Our model achieves state-of-the-art results on the Breakfast, Assembly101 and 50Salads datasets in both stochastic and deterministic settings. Code: https://github.com/olga-zats/GTDA .
title Gated Temporal Diffusion for Stochastic Long-Term Dense Anticipation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.11954