Ego-VPA: Egocentric Video Understanding with Parameter-efficient Adaptation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915174384926720 |
|---|---|
| author | Wu, Tz-Ying Min, Kyle Tripathi, Subarna Vasconcelos, Nuno |
| author_facet | Wu, Tz-Ying Min, Kyle Tripathi, Subarna Vasconcelos, Nuno |
| contents | Video understanding typically requires fine-tuning the large backbone when adapting to new domains. In this paper, we leverage the egocentric video foundation models (Ego-VFMs) based on video-language pre-training and propose a parameter-efficient adaptation for egocentric video tasks, namely Ego-VPA. It employs a local sparse approximation for each video frame/text feature using the basis prompts, and the selected basis prompts are used to synthesize video/text prompts. Since the basis prompts are shared across frames and modalities, it models context fusion and cross-modal transfer in an efficient fashion. Experiments show that Ego-VPA excels in lightweight adaptation (with only 0.84% learnable parameters), largely improving over baselines and reaching the performance of full fine-tuning. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_19520 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Ego-VPA: Egocentric Video Understanding with Parameter-efficient Adaptation Wu, Tz-Ying Min, Kyle Tripathi, Subarna Vasconcelos, Nuno Computer Vision and Pattern Recognition Machine Learning Video understanding typically requires fine-tuning the large backbone when adapting to new domains. In this paper, we leverage the egocentric video foundation models (Ego-VFMs) based on video-language pre-training and propose a parameter-efficient adaptation for egocentric video tasks, namely Ego-VPA. It employs a local sparse approximation for each video frame/text feature using the basis prompts, and the selected basis prompts are used to synthesize video/text prompts. Since the basis prompts are shared across frames and modalities, it models context fusion and cross-modal transfer in an efficient fashion. Experiments show that Ego-VPA excels in lightweight adaptation (with only 0.84% learnable parameters), largely improving over baselines and reaching the performance of full fine-tuning. |
| title | Ego-VPA: Egocentric Video Understanding with Parameter-efficient Adaptation |
| topic | Computer Vision and Pattern Recognition Machine Learning |
| url | https://arxiv.org/abs/2407.19520 |