Active Multimodal Distillation for Few-shot Action Recognition

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Feng, Weijia, Zhu, Yichen, Zhang, Ruojia, Wang, Chenyang, Ma, Fei, Wang, Xiaobao, Li, Xiaobai
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908408790122496
author Feng, Weijia
Zhu, Yichen
Zhang, Ruojia
Wang, Chenyang
Ma, Fei
Wang, Xiaobao
Li, Xiaobai
author_facet Feng, Weijia
Zhu, Yichen
Zhang, Ruojia
Wang, Chenyang
Ma, Fei
Wang, Xiaobao
Li, Xiaobai
contents Owing to its rapid progress and broad application prospects, few-shot action recognition has attracted considerable interest. However, current methods are predominantly based on limited single-modal data, which does not fully exploit the potential of multimodal information. This paper presents a novel framework that actively identifies reliable modalities for each sample using task-specific contextual cues, thus significantly improving recognition performance. Our framework integrates an Active Sample Inference (ASI) module, which utilizes active inference to predict reliable modalities based on posterior distributions and subsequently organizes them accordingly. Unlike reinforcement learning, active inference replaces rewards with evidence-based preferences, making more stable predictions. Additionally, we introduce an active mutual distillation module that enhances the representation learning of less reliable modalities by transferring knowledge from more reliable ones. Adaptive multimodal inference is employed during the meta-test to assign higher weights to reliable modalities. Extensive experiments across multiple benchmarks demonstrate that our method significantly outperforms existing approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13322
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Active Multimodal Distillation for Few-shot Action Recognition
Feng, Weijia
Zhu, Yichen
Zhang, Ruojia
Wang, Chenyang
Ma, Fei
Wang, Xiaobao
Li, Xiaobai
Computer Vision and Pattern Recognition
Artificial Intelligence
Owing to its rapid progress and broad application prospects, few-shot action recognition has attracted considerable interest. However, current methods are predominantly based on limited single-modal data, which does not fully exploit the potential of multimodal information. This paper presents a novel framework that actively identifies reliable modalities for each sample using task-specific contextual cues, thus significantly improving recognition performance. Our framework integrates an Active Sample Inference (ASI) module, which utilizes active inference to predict reliable modalities based on posterior distributions and subsequently organizes them accordingly. Unlike reinforcement learning, active inference replaces rewards with evidence-based preferences, making more stable predictions. Additionally, we introduce an active mutual distillation module that enhances the representation learning of less reliable modalities by transferring knowledge from more reliable ones. Adaptive multimodal inference is employed during the meta-test to assign higher weights to reliable modalities. Extensive experiments across multiple benchmarks demonstrate that our method significantly outperforms existing approaches.
title Active Multimodal Distillation for Few-shot Action Recognition
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.13322