M2-CLIP: A Multimodal, Multi-task Adapting Framework for Video Action Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Mengmeng, Xing, Jiazheng, Jiang, Boyuan, Chen, Jun, Mei, Jianbiao, Zuo, Xingxing, Dai, Guang, Wang, Jingdong, Liu, Yong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911762182307840
author Wang, Mengmeng
Xing, Jiazheng
Jiang, Boyuan
Chen, Jun
Mei, Jianbiao
Zuo, Xingxing
Dai, Guang
Wang, Jingdong
Liu, Yong
author_facet Wang, Mengmeng
Xing, Jiazheng
Jiang, Boyuan
Chen, Jun
Mei, Jianbiao
Zuo, Xingxing
Dai, Guang
Wang, Jingdong
Liu, Yong
contents Recently, the rise of large-scale vision-language pretrained models like CLIP, coupled with the technology of Parameter-Efficient FineTuning (PEFT), has captured substantial attraction in video action recognition. Nevertheless, prevailing approaches tend to prioritize strong supervised performance at the expense of compromising the models' generalization capabilities during transfer. In this paper, we introduce a novel Multimodal, Multi-task CLIP adapting framework named \name to address these challenges, preserving both high supervised performance and robust transferability. Firstly, to enhance the individual modality architectures, we introduce multimodal adapters to both the visual and text branches. Specifically, we design a novel visual TED-Adapter, that performs global Temporal Enhancement and local temporal Difference modeling to improve the temporal representation capabilities of the visual encoder. Moreover, we adopt text encoder adapters to strengthen the learning of semantic label information. Secondly, we design a multi-task decoder with a rich set of supervisory signals to adeptly satisfy the need for strong supervised performance and generalization within a multimodal framework. Experimental results validate the efficacy of our approach, demonstrating exceptional performance in supervised learning while maintaining strong generalization in zero-shot scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2401_11649
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle M2-CLIP: A Multimodal, Multi-task Adapting Framework for Video Action Recognition
Wang, Mengmeng
Xing, Jiazheng
Jiang, Boyuan
Chen, Jun
Mei, Jianbiao
Zuo, Xingxing
Dai, Guang
Wang, Jingdong
Liu, Yong
Computer Vision and Pattern Recognition
Recently, the rise of large-scale vision-language pretrained models like CLIP, coupled with the technology of Parameter-Efficient FineTuning (PEFT), has captured substantial attraction in video action recognition. Nevertheless, prevailing approaches tend to prioritize strong supervised performance at the expense of compromising the models' generalization capabilities during transfer. In this paper, we introduce a novel Multimodal, Multi-task CLIP adapting framework named \name to address these challenges, preserving both high supervised performance and robust transferability. Firstly, to enhance the individual modality architectures, we introduce multimodal adapters to both the visual and text branches. Specifically, we design a novel visual TED-Adapter, that performs global Temporal Enhancement and local temporal Difference modeling to improve the temporal representation capabilities of the visual encoder. Moreover, we adopt text encoder adapters to strengthen the learning of semantic label information. Secondly, we design a multi-task decoder with a rich set of supervisory signals to adeptly satisfy the need for strong supervised performance and generalization within a multimodal framework. Experimental results validate the efficacy of our approach, demonstrating exceptional performance in supervised learning while maintaining strong generalization in zero-shot scenarios.
title M2-CLIP: A Multimodal, Multi-task Adapting Framework for Video Action Recognition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.11649