Gaze-Guided Graph Neural Network for Action Anticipation Conditioned on Intention
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909167688613888 |
|---|---|
| author | Ozdel, Suleyman Rong, Yao Albaba, Berat Mert Kuo, Yen-Ling Wang, Xi Kasneci, Enkelejda |
| author_facet | Ozdel, Suleyman Rong, Yao Albaba, Berat Mert Kuo, Yen-Ling Wang, Xi Kasneci, Enkelejda |
| contents | Humans utilize their gaze to concentrate on essential information while perceiving and interpreting intentions in videos. Incorporating human gaze into computational algorithms can significantly enhance model performance in video understanding tasks. In this work, we address a challenging and innovative task in video understanding: predicting the actions of an agent in a video based on a partial video. We introduce the Gaze-guided Action Anticipation algorithm, which establishes a visual-semantic graph from the video input. Our method utilizes a Graph Neural Network to recognize the agent's intention and predict the action sequence to fulfill this intention. To assess the efficiency of our approach, we collect a dataset containing household activities generated in the VirtualHome environment, accompanied by human gaze data of viewing videos. Our method outperforms state-of-the-art techniques, achieving a 7\% improvement in accuracy for 18-class intention recognition. This highlights the efficiency of our method in learning important features from human gaze data. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2404_07347 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Gaze-Guided Graph Neural Network for Action Anticipation Conditioned on Intention Ozdel, Suleyman Rong, Yao Albaba, Berat Mert Kuo, Yen-Ling Wang, Xi Kasneci, Enkelejda Computer Vision and Pattern Recognition Human-Computer Interaction Machine Learning Humans utilize their gaze to concentrate on essential information while perceiving and interpreting intentions in videos. Incorporating human gaze into computational algorithms can significantly enhance model performance in video understanding tasks. In this work, we address a challenging and innovative task in video understanding: predicting the actions of an agent in a video based on a partial video. We introduce the Gaze-guided Action Anticipation algorithm, which establishes a visual-semantic graph from the video input. Our method utilizes a Graph Neural Network to recognize the agent's intention and predict the action sequence to fulfill this intention. To assess the efficiency of our approach, we collect a dataset containing household activities generated in the VirtualHome environment, accompanied by human gaze data of viewing videos. Our method outperforms state-of-the-art techniques, achieving a 7\% improvement in accuracy for 18-class intention recognition. This highlights the efficiency of our method in learning important features from human gaze data. |
| title | Gaze-Guided Graph Neural Network for Action Anticipation Conditioned on Intention |
| topic | Computer Vision and Pattern Recognition Human-Computer Interaction Machine Learning |
| url | https://arxiv.org/abs/2404.07347 |