Ego-VPA: Egocentric Video Understanding with Parameter-efficient Adaptation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Tz-Ying, Min, Kyle, Tripathi, Subarna, Vasconcelos, Nuno
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915174384926720
author Wu, Tz-Ying
Min, Kyle
Tripathi, Subarna
Vasconcelos, Nuno
author_facet Wu, Tz-Ying
Min, Kyle
Tripathi, Subarna
Vasconcelos, Nuno
contents Video understanding typically requires fine-tuning the large backbone when adapting to new domains. In this paper, we leverage the egocentric video foundation models (Ego-VFMs) based on video-language pre-training and propose a parameter-efficient adaptation for egocentric video tasks, namely Ego-VPA. It employs a local sparse approximation for each video frame/text feature using the basis prompts, and the selected basis prompts are used to synthesize video/text prompts. Since the basis prompts are shared across frames and modalities, it models context fusion and cross-modal transfer in an efficient fashion. Experiments show that Ego-VPA excels in lightweight adaptation (with only 0.84% learnable parameters), largely improving over baselines and reaching the performance of full fine-tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2407_19520
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Ego-VPA: Egocentric Video Understanding with Parameter-efficient Adaptation
Wu, Tz-Ying
Min, Kyle
Tripathi, Subarna
Vasconcelos, Nuno
Computer Vision and Pattern Recognition
Machine Learning
Video understanding typically requires fine-tuning the large backbone when adapting to new domains. In this paper, we leverage the egocentric video foundation models (Ego-VFMs) based on video-language pre-training and propose a parameter-efficient adaptation for egocentric video tasks, namely Ego-VPA. It employs a local sparse approximation for each video frame/text feature using the basis prompts, and the selected basis prompts are used to synthesize video/text prompts. Since the basis prompts are shared across frames and modalities, it models context fusion and cross-modal transfer in an efficient fashion. Experiments show that Ego-VPA excels in lightweight adaptation (with only 0.84% learnable parameters), largely improving over baselines and reaching the performance of full fine-tuning.
title Ego-VPA: Egocentric Video Understanding with Parameter-efficient Adaptation
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2407.19520