Gaze-Guided Graph Neural Network for Action Anticipation Conditioned on Intention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ozdel, Suleyman, Rong, Yao, Albaba, Berat Mert, Kuo, Yen-Ling, Wang, Xi, Kasneci, Enkelejda
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909167688613888
author Ozdel, Suleyman
Rong, Yao
Albaba, Berat Mert
Kuo, Yen-Ling
Wang, Xi
Kasneci, Enkelejda
author_facet Ozdel, Suleyman
Rong, Yao
Albaba, Berat Mert
Kuo, Yen-Ling
Wang, Xi
Kasneci, Enkelejda
contents Humans utilize their gaze to concentrate on essential information while perceiving and interpreting intentions in videos. Incorporating human gaze into computational algorithms can significantly enhance model performance in video understanding tasks. In this work, we address a challenging and innovative task in video understanding: predicting the actions of an agent in a video based on a partial video. We introduce the Gaze-guided Action Anticipation algorithm, which establishes a visual-semantic graph from the video input. Our method utilizes a Graph Neural Network to recognize the agent's intention and predict the action sequence to fulfill this intention. To assess the efficiency of our approach, we collect a dataset containing household activities generated in the VirtualHome environment, accompanied by human gaze data of viewing videos. Our method outperforms state-of-the-art techniques, achieving a 7\% improvement in accuracy for 18-class intention recognition. This highlights the efficiency of our method in learning important features from human gaze data.
format Preprint
id arxiv_https___arxiv_org_abs_2404_07347
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Gaze-Guided Graph Neural Network for Action Anticipation Conditioned on Intention
Ozdel, Suleyman
Rong, Yao
Albaba, Berat Mert
Kuo, Yen-Ling
Wang, Xi
Kasneci, Enkelejda
Computer Vision and Pattern Recognition
Human-Computer Interaction
Machine Learning
Humans utilize their gaze to concentrate on essential information while perceiving and interpreting intentions in videos. Incorporating human gaze into computational algorithms can significantly enhance model performance in video understanding tasks. In this work, we address a challenging and innovative task in video understanding: predicting the actions of an agent in a video based on a partial video. We introduce the Gaze-guided Action Anticipation algorithm, which establishes a visual-semantic graph from the video input. Our method utilizes a Graph Neural Network to recognize the agent's intention and predict the action sequence to fulfill this intention. To assess the efficiency of our approach, we collect a dataset containing household activities generated in the VirtualHome environment, accompanied by human gaze data of viewing videos. Our method outperforms state-of-the-art techniques, achieving a 7\% improvement in accuracy for 18-class intention recognition. This highlights the efficiency of our method in learning important features from human gaze data.
title Gaze-Guided Graph Neural Network for Action Anticipation Conditioned on Intention
topic Computer Vision and Pattern Recognition
Human-Computer Interaction
Machine Learning
url https://arxiv.org/abs/2404.07347