Efficient Egocentric Action Recognition with Multimodal Data

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Calzavara, Marco, Kastrati, Ard, Macchini, Matteo, Vasilevski, Dushan, Wattenhofer, Roger
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916772953718784
author Calzavara, Marco
Kastrati, Ard
Macchini, Matteo
Vasilevski, Dushan
Wattenhofer, Roger
author_facet Calzavara, Marco
Kastrati, Ard
Macchini, Matteo
Vasilevski, Dushan
Wattenhofer, Roger
contents The increasing availability of wearable XR devices opens new perspectives for Egocentric Action Recognition (EAR) systems, which can provide deeper human understanding and situation awareness. However, deploying real-time algorithms on these devices can be challenging due to the inherent trade-offs between portability, battery life, and computational resources. In this work, we systematically analyze the impact of sampling frequency across different input modalities - RGB video and 3D hand pose - on egocentric action recognition performance and CPU usage. By exploring a range of configurations, we provide a comprehensive characterization of the trade-offs between accuracy and computational efficiency. Our findings reveal that reducing the sampling rate of RGB frames, when complemented with higher-frequency 3D hand pose input, can preserve high accuracy while significantly lowering CPU demands. Notably, we observe up to a 3x reduction in CPU usage with minimal to no loss in recognition performance. This highlights the potential of multimodal input strategies as a viable approach to achieving efficient, real-time EAR on XR devices.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01757
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Egocentric Action Recognition with Multimodal Data
Calzavara, Marco
Kastrati, Ard
Macchini, Matteo
Vasilevski, Dushan
Wattenhofer, Roger
Computer Vision and Pattern Recognition
Artificial Intelligence
The increasing availability of wearable XR devices opens new perspectives for Egocentric Action Recognition (EAR) systems, which can provide deeper human understanding and situation awareness. However, deploying real-time algorithms on these devices can be challenging due to the inherent trade-offs between portability, battery life, and computational resources. In this work, we systematically analyze the impact of sampling frequency across different input modalities - RGB video and 3D hand pose - on egocentric action recognition performance and CPU usage. By exploring a range of configurations, we provide a comprehensive characterization of the trade-offs between accuracy and computational efficiency. Our findings reveal that reducing the sampling rate of RGB frames, when complemented with higher-frequency 3D hand pose input, can preserve high accuracy while significantly lowering CPU demands. Notably, we observe up to a 3x reduction in CPU usage with minimal to no loss in recognition performance. This highlights the potential of multimodal input strategies as a viable approach to achieving efficient, real-time EAR on XR devices.
title Efficient Egocentric Action Recognition with Multimodal Data
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.01757