Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2503.12678 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915200460914688 |
|---|---|
| author | Ghosh, Partho Hossain, Raisa Bentay Zunaed, Mohammad Hasan, Taufiq |
| author_facet | Ghosh, Partho Hossain, Raisa Bentay Zunaed, Mohammad Hasan, Taufiq |
| contents | Automatic video activity recognition is crucial across numerous domains like surveillance, healthcare, and robotics. However, recognizing human activities from video data becomes challenging when training and test data stem from diverse domains. Domain generalization, adapting to unforeseen domains, is thus essential. This paper focuses on office activity recognition amidst environmental variability. We propose three pre-processing techniques applicable to any video encoder, enhancing robustness against environmental variations. Our study showcases the efficacy of MViT, a leading state-of-the-art video classification model, and other video encoders combined with our techniques, outperforming state-of-the-art domain adaptation methods. Our approach significantly boosts accuracy, precision, recall and F1 score on unseen domains, emphasizing its adaptability in real-world scenarios with diverse video data sources. This method lays a foundation for more reliable video activity recognition systems across heterogeneous data domains. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_12678 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Domain Generalization for Improved Human Activity Recognition in Office Space Videos Using Adaptive Pre-processing Ghosh, Partho Hossain, Raisa Bentay Zunaed, Mohammad Hasan, Taufiq Computer Vision and Pattern Recognition Automatic video activity recognition is crucial across numerous domains like surveillance, healthcare, and robotics. However, recognizing human activities from video data becomes challenging when training and test data stem from diverse domains. Domain generalization, adapting to unforeseen domains, is thus essential. This paper focuses on office activity recognition amidst environmental variability. We propose three pre-processing techniques applicable to any video encoder, enhancing robustness against environmental variations. Our study showcases the efficacy of MViT, a leading state-of-the-art video classification model, and other video encoders combined with our techniques, outperforming state-of-the-art domain adaptation methods. Our approach significantly boosts accuracy, precision, recall and F1 score on unseen domains, emphasizing its adaptability in real-world scenarios with diverse video data sources. This method lays a foundation for more reliable video activity recognition systems across heterogeneous data domains. |
| title | Domain Generalization for Improved Human Activity Recognition in Office Space Videos Using Adaptive Pre-processing |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2503.12678 |