Leveraging Vision-Language Pre-training for Human Activity Recognition in Still Images

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mahanta, Cristina, Bhatia, Gagan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915346680643584
author Mahanta, Cristina
Bhatia, Gagan
author_facet Mahanta, Cristina
Bhatia, Gagan
contents Recognising human activity in a single photo enables indexing, safety and assistive applications, yet lacks motion cues. Using 285 MSCOCO images labelled as walking, running, sitting, and standing, scratch CNNs scored 41% accuracy. Fine-tuning multimodal CLIP raised this to 76%, demonstrating that contrastive vision-language pre-training decisively improves still-image action recognition in real-world deployments.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13458
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Leveraging Vision-Language Pre-training for Human Activity Recognition in Still Images
Mahanta, Cristina
Bhatia, Gagan
Computer Vision and Pattern Recognition
Computation and Language
Recognising human activity in a single photo enables indexing, safety and assistive applications, yet lacks motion cues. Using 285 MSCOCO images labelled as walking, running, sitting, and standing, scratch CNNs scored 41% accuracy. Fine-tuning multimodal CLIP raised this to 76%, demonstrating that contrastive vision-language pre-training decisively improves still-image action recognition in real-world deployments.
title Leveraging Vision-Language Pre-training for Human Activity Recognition in Still Images
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2506.13458