Keystep Recognition using Graph Neural Networks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Romero, Julia Lee, Min, Kyle, Tripathi, Subarna, Karimzadeh, Morteza
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908823884660736
author Romero, Julia Lee
Min, Kyle
Tripathi, Subarna
Karimzadeh, Morteza
author_facet Romero, Julia Lee
Min, Kyle
Tripathi, Subarna
Karimzadeh, Morteza
contents We pose keystep recognition as a node classification task, and propose a flexible graph-learning framework for fine-grained keystep recognition that is able to effectively leverage long-term dependencies in egocentric videos. Our approach, termed GLEVR, consists of constructing a graph where each video clip of the egocentric video corresponds to a node. The constructed graphs are sparse and computationally efficient, outperforming existing larger models substantially. We further leverage alignment between egocentric and exocentric videos during training for improved inference on egocentric videos, as well as adding automatic captioning as an additional modality. We consider each clip of each exocentric video (if available) or video captions as additional nodes during training. We examine several strategies to define connections across these nodes. We perform extensive experiments on the Ego-Exo4D dataset and show that our proposed flexible graph-based framework notably outperforms existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01102
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Keystep Recognition using Graph Neural Networks
Romero, Julia Lee
Min, Kyle
Tripathi, Subarna
Karimzadeh, Morteza
Computer Vision and Pattern Recognition
We pose keystep recognition as a node classification task, and propose a flexible graph-learning framework for fine-grained keystep recognition that is able to effectively leverage long-term dependencies in egocentric videos. Our approach, termed GLEVR, consists of constructing a graph where each video clip of the egocentric video corresponds to a node. The constructed graphs are sparse and computationally efficient, outperforming existing larger models substantially. We further leverage alignment between egocentric and exocentric videos during training for improved inference on egocentric videos, as well as adding automatic captioning as an additional modality. We consider each clip of each exocentric video (if available) or video captions as additional nodes during training. We examine several strategies to define connections across these nodes. We perform extensive experiments on the Ego-Exo4D dataset and show that our proposed flexible graph-based framework notably outperforms existing methods.
title Keystep Recognition using Graph Neural Networks
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.01102