Beyond Label Semantics: Language-Guided Action Anatomy for Few-shot Action Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qian, Zefeng, Yao, Xincheng, Huang, Yifei, Zhang, Chongyang, Ying, Jiangyong, Sun, Hong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908499878871040
author Qian, Zefeng
Yao, Xincheng
Huang, Yifei
Zhang, Chongyang
Ying, Jiangyong
Sun, Hong
author_facet Qian, Zefeng
Yao, Xincheng
Huang, Yifei
Zhang, Chongyang
Ying, Jiangyong
Sun, Hong
contents Few-shot action recognition (FSAR) aims to classify human actions in videos with only a small number of labeled samples per category. The scarcity of training data has driven recent efforts to incorporate additional modalities, particularly text. However, the subtle variations in human posture, motion dynamics, and the object interactions that occur during different phases, are critical inherent knowledge of actions that cannot be fully exploited by action labels alone. In this work, we propose Language-Guided Action Anatomy (LGA), a novel framework that goes beyond label semantics by leveraging Large Language Models (LLMs) to dissect the essential representational characteristics hidden beneath action labels. Guided by the prior knowledge encoded in LLM, LGA effectively captures rich spatiotemporal cues in few-shot scenarios. Specifically, for text, we prompt an off-the-shelf LLM to anatomize labels into sequences of atomic action descriptions, focusing on the three core elements of action (subject, motion, object). For videos, a Visual Anatomy Module segments actions into atomic video phases to capture the sequential structure of actions. A fine-grained fusion strategy then integrates textual and visual features at the atomic level, resulting in more generalizable prototypes. Finally, we introduce a Multimodal Matching mechanism, comprising both video-video and video-text matching, to ensure robust few-shot classification. Experimental results demonstrate that LGA achieves state-of-the-art performance across multipe FSAR benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2507_16287
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Label Semantics: Language-Guided Action Anatomy for Few-shot Action Recognition
Qian, Zefeng
Yao, Xincheng
Huang, Yifei
Zhang, Chongyang
Ying, Jiangyong
Sun, Hong
Computer Vision and Pattern Recognition
Few-shot action recognition (FSAR) aims to classify human actions in videos with only a small number of labeled samples per category. The scarcity of training data has driven recent efforts to incorporate additional modalities, particularly text. However, the subtle variations in human posture, motion dynamics, and the object interactions that occur during different phases, are critical inherent knowledge of actions that cannot be fully exploited by action labels alone. In this work, we propose Language-Guided Action Anatomy (LGA), a novel framework that goes beyond label semantics by leveraging Large Language Models (LLMs) to dissect the essential representational characteristics hidden beneath action labels. Guided by the prior knowledge encoded in LLM, LGA effectively captures rich spatiotemporal cues in few-shot scenarios. Specifically, for text, we prompt an off-the-shelf LLM to anatomize labels into sequences of atomic action descriptions, focusing on the three core elements of action (subject, motion, object). For videos, a Visual Anatomy Module segments actions into atomic video phases to capture the sequential structure of actions. A fine-grained fusion strategy then integrates textual and visual features at the atomic level, resulting in more generalizable prototypes. Finally, we introduce a Multimodal Matching mechanism, comprising both video-video and video-text matching, to ensure robust few-shot classification. Experimental results demonstrate that LGA achieves state-of-the-art performance across multipe FSAR benchmarks.
title Beyond Label Semantics: Language-Guided Action Anatomy for Few-shot Action Recognition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.16287