Group Relative Augmentation for Data Efficient Action Detection

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Patel, Deep Anil, Melvin, Iain, Izzo, Zachary, Min, Martin Renqiang
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909710109638656
author Patel, Deep Anil
Melvin, Iain
Izzo, Zachary
Min, Martin Renqiang
author_facet Patel, Deep Anil
Melvin, Iain
Izzo, Zachary
Min, Martin Renqiang
contents Adapting large Video-Language Models (VLMs) for action detection using only a few examples poses challenges like overfitting and the granularity mismatch between scene-level pre-training and required person-centric understanding. We propose an efficient adaptation strategy combining parameter-efficient tuning (LoRA) with a novel learnable internal feature augmentation. Applied within the frozen VLM backbone using FiLM, these augmentations generate diverse feature variations directly relevant to the task. Additionally, we introduce a group-weighted loss function that dynamically modulates the training contribution of each augmented sample based on its prediction divergence relative to the group average. This promotes robust learning by prioritizing informative yet reasonable augmentations. We demonstrate our method's effectiveness on complex multi-label, multi-person action detection datasets (AVA, MOMA), achieving strong mAP performance and showcasing significant data efficiency for adapting VLMs from limited examples.
format Preprint
id arxiv_https___arxiv_org_abs_2507_21353
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Group Relative Augmentation for Data Efficient Action Detection
Patel, Deep Anil
Melvin, Iain
Izzo, Zachary
Min, Martin Renqiang
Computer Vision and Pattern Recognition
Machine Learning
Adapting large Video-Language Models (VLMs) for action detection using only a few examples poses challenges like overfitting and the granularity mismatch between scene-level pre-training and required person-centric understanding. We propose an efficient adaptation strategy combining parameter-efficient tuning (LoRA) with a novel learnable internal feature augmentation. Applied within the frozen VLM backbone using FiLM, these augmentations generate diverse feature variations directly relevant to the task. Additionally, we introduce a group-weighted loss function that dynamically modulates the training contribution of each augmented sample based on its prediction divergence relative to the group average. This promotes robust learning by prioritizing informative yet reasonable augmentations. We demonstrate our method's effectiveness on complex multi-label, multi-person action detection datasets (AVA, MOMA), achieving strong mAP performance and showcasing significant data efficiency for adapting VLMs from limited examples.
title Group Relative Augmentation for Data Efficient Action Detection
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2507.21353