Computer Vision based group activity detection and action spotting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sivalingam, Narthana, Sivasthigan, Santhirarajah, Mahendranathan, Thamayanthi, Godaliyadda, G. M. R. I., Ekanayake, M. P. B., Herath, H. M. V. R.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908659446972416
author Sivalingam, Narthana
Sivasthigan, Santhirarajah
Mahendranathan, Thamayanthi
Godaliyadda, G. M. R. I.
Ekanayake, M. P. B.
Herath, H. M. V. R.
author_facet Sivalingam, Narthana
Sivasthigan, Santhirarajah
Mahendranathan, Thamayanthi
Godaliyadda, G. M. R. I.
Ekanayake, M. P. B.
Herath, H. M. V. R.
contents Group activity detection in multi-person scenes is challenging due to complex human interactions, occlusions, and variations in appearance over time. This work presents a computer vision based framework for group activity recognition and action spotting using a combination of deep learning models and graph based relational reasoning. The system first applies Mask R-CNN to obtain accurate actor localization through bounding boxes and instance masks. Multiple backbone networks, including Inception V3, MobileNet, and VGG16, are used to extract feature maps, and RoIAlign is applied to preserve spatial alignment when generating actor specific features. The mask information is then fused with the feature maps to obtain refined masked feature representations for each actor. To model interactions between individuals, we construct Actor Relation Graphs that encode appearance similarity and positional relations using methods such as normalized cross correlation, sum of absolute differences, and dot product. Graph Convolutional Networks operate on these graphs to reason about relationships and predict both individual actions and group level activities. Experiments on the Collective Activity dataset demonstrate that the combination of mask based feature refinement, robust similarity search, and graph neural network reasoning leads to improved recognition performance across both crowded and non crowded scenarios. This approach highlights the potential of integrating segmentation, feature extraction, and relational graph reasoning for complex video understanding tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2511_13315
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Computer Vision based group activity detection and action spotting
Sivalingam, Narthana
Sivasthigan, Santhirarajah
Mahendranathan, Thamayanthi
Godaliyadda, G. M. R. I.
Ekanayake, M. P. B.
Herath, H. M. V. R.
Computer Vision and Pattern Recognition
Artificial Intelligence
Group activity detection in multi-person scenes is challenging due to complex human interactions, occlusions, and variations in appearance over time. This work presents a computer vision based framework for group activity recognition and action spotting using a combination of deep learning models and graph based relational reasoning. The system first applies Mask R-CNN to obtain accurate actor localization through bounding boxes and instance masks. Multiple backbone networks, including Inception V3, MobileNet, and VGG16, are used to extract feature maps, and RoIAlign is applied to preserve spatial alignment when generating actor specific features. The mask information is then fused with the feature maps to obtain refined masked feature representations for each actor. To model interactions between individuals, we construct Actor Relation Graphs that encode appearance similarity and positional relations using methods such as normalized cross correlation, sum of absolute differences, and dot product. Graph Convolutional Networks operate on these graphs to reason about relationships and predict both individual actions and group level activities. Experiments on the Collective Activity dataset demonstrate that the combination of mask based feature refinement, robust similarity search, and graph neural network reasoning leads to improved recognition performance across both crowded and non crowded scenarios. This approach highlights the potential of integrating segmentation, feature extraction, and relational graph reasoning for complex video understanding tasks.
title Computer Vision based group activity detection and action spotting
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2511.13315