Human Action Co-occurrence in Lifestyle Vlogs using Graph Link Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ignat, Oana, Castro, Santiago, Li, Weiji, Mihalcea, Rada
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914839554686976
author Ignat, Oana
Castro, Santiago
Li, Weiji
Mihalcea, Rada
author_facet Ignat, Oana
Castro, Santiago
Li, Weiji
Mihalcea, Rada
contents We introduce the task of automatic human action co-occurrence identification, i.e., determine whether two human actions can co-occur in the same interval of time. We create and make publicly available the ACE (Action Co-occurrencE) dataset, consisting of a large graph of ~12k co-occurring pairs of visual actions and their corresponding video clips. We describe graph link prediction models that leverage visual and textual information to automatically infer if two actions are co-occurring. We show that graphs are particularly well suited to capture relations between human actions, and the learned graph representations are effective for our task and capture novel and relevant information across different data domains. The ACE dataset and the code introduced in this paper are publicly available at https://github.com/MichiganNLP/vlog_action_co-occurrence.
format Preprint
id arxiv_https___arxiv_org_abs_2309_06219
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Human Action Co-occurrence in Lifestyle Vlogs using Graph Link Prediction
Ignat, Oana
Castro, Santiago
Li, Weiji
Mihalcea, Rada
Computer Vision and Pattern Recognition
Computation and Language
Computers and Society
Information Retrieval
We introduce the task of automatic human action co-occurrence identification, i.e., determine whether two human actions can co-occur in the same interval of time. We create and make publicly available the ACE (Action Co-occurrencE) dataset, consisting of a large graph of ~12k co-occurring pairs of visual actions and their corresponding video clips. We describe graph link prediction models that leverage visual and textual information to automatically infer if two actions are co-occurring. We show that graphs are particularly well suited to capture relations between human actions, and the learned graph representations are effective for our task and capture novel and relevant information across different data domains. The ACE dataset and the code introduced in this paper are publicly available at https://github.com/MichiganNLP/vlog_action_co-occurrence.
title Human Action Co-occurrence in Lifestyle Vlogs using Graph Link Prediction
topic Computer Vision and Pattern Recognition
Computation and Language
Computers and Society
Information Retrieval
url https://arxiv.org/abs/2309.06219