X-MIC: Cross-Modal Instance Conditioning for Egocentric Action Generalization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kukleva, Anna, Sener, Fadime, Remelli, Edoardo, Tekin, Bugra, Sauser, Eric, Schiele, Bernt, Ma, Shugao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910389814427648
author Kukleva, Anna
Sener, Fadime
Remelli, Edoardo
Tekin, Bugra
Sauser, Eric
Schiele, Bernt
Ma, Shugao
author_facet Kukleva, Anna
Sener, Fadime
Remelli, Edoardo
Tekin, Bugra
Sauser, Eric
Schiele, Bernt
Ma, Shugao
contents Lately, there has been growing interest in adapting vision-language models (VLMs) to image and third-person video classification due to their success in zero-shot recognition. However, the adaptation of these models to egocentric videos has been largely unexplored. To address this gap, we propose a simple yet effective cross-modal adaptation framework, which we call X-MIC. Using a video adapter, our pipeline learns to align frozen text embeddings to each egocentric video directly in the shared embedding space. Our novel adapter architecture retains and improves generalization of the pre-trained VLMs by disentangling learnable temporal modeling and frozen visual encoder. This results in an enhanced alignment of text embeddings to each egocentric video, leading to a significant improvement in cross-dataset generalization. We evaluate our approach on the Epic-Kitchens, Ego4D, and EGTEA datasets for fine-grained cross-dataset action generalization, demonstrating the effectiveness of our method. Code is available at https://github.com/annusha/xmic
format Preprint
id arxiv_https___arxiv_org_abs_2403_19811
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle X-MIC: Cross-Modal Instance Conditioning for Egocentric Action Generalization
Kukleva, Anna
Sener, Fadime
Remelli, Edoardo
Tekin, Bugra
Sauser, Eric
Schiele, Bernt
Ma, Shugao
Computer Vision and Pattern Recognition
Lately, there has been growing interest in adapting vision-language models (VLMs) to image and third-person video classification due to their success in zero-shot recognition. However, the adaptation of these models to egocentric videos has been largely unexplored. To address this gap, we propose a simple yet effective cross-modal adaptation framework, which we call X-MIC. Using a video adapter, our pipeline learns to align frozen text embeddings to each egocentric video directly in the shared embedding space. Our novel adapter architecture retains and improves generalization of the pre-trained VLMs by disentangling learnable temporal modeling and frozen visual encoder. This results in an enhanced alignment of text embeddings to each egocentric video, leading to a significant improvement in cross-dataset generalization. We evaluate our approach on the Epic-Kitchens, Ego4D, and EGTEA datasets for fine-grained cross-dataset action generalization, demonstrating the effectiveness of our method. Code is available at https://github.com/annusha/xmic
title X-MIC: Cross-Modal Instance Conditioning for Egocentric Action Generalization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.19811