MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Yuhuan, Ma, Chaofan, Mao, Zhenjie, Yao, Jiangchao, Zhang, Ya, Wang, Yanfeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918075172913152
author Yang, Yuhuan
Ma, Chaofan
Mao, Zhenjie
Yao, Jiangchao
Zhang, Ya
Wang, Yanfeng
author_facet Yang, Yuhuan
Ma, Chaofan
Mao, Zhenjie
Yao, Jiangchao
Zhang, Ya
Wang, Yanfeng
contents Video understanding is a complex challenge that requires effective modeling of spatial-temporal dynamics. With the success of image foundation models (IFMs) in image understanding, recent approaches have explored parameter-efficient fine-tuning (PEFT) to adapt IFMs for video. However, most of these methods tend to process spatial and temporal information separately, which may fail to capture the full intricacy of video dynamics. In this paper, we propose MoMa, an efficient adapter framework that achieves full spatial-temporal modeling by integrating Mamba's selective state space modeling into IFMs. We propose a novel SeqMod operation to inject spatial-temporal information into pre-trained IFMs, without disrupting their original features. By incorporating SeqMod into a Divide-and-Modulate architecture, MoMa enhances video understanding while maintaining computational efficiency. Extensive experiments on multiple video benchmarks demonstrate the effectiveness of MoMa, achieving superior performance with reduced computational cost.
format Preprint
id arxiv_https___arxiv_org_abs_2506_23283
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition
Yang, Yuhuan
Ma, Chaofan
Mao, Zhenjie
Yao, Jiangchao
Zhang, Ya
Wang, Yanfeng
Computer Vision and Pattern Recognition
Video understanding is a complex challenge that requires effective modeling of spatial-temporal dynamics. With the success of image foundation models (IFMs) in image understanding, recent approaches have explored parameter-efficient fine-tuning (PEFT) to adapt IFMs for video. However, most of these methods tend to process spatial and temporal information separately, which may fail to capture the full intricacy of video dynamics. In this paper, we propose MoMa, an efficient adapter framework that achieves full spatial-temporal modeling by integrating Mamba's selective state space modeling into IFMs. We propose a novel SeqMod operation to inject spatial-temporal information into pre-trained IFMs, without disrupting their original features. By incorporating SeqMod into a Divide-and-Modulate architecture, MoMa enhances video understanding while maintaining computational efficiency. Extensive experiments on multiple video benchmarks demonstrate the effectiveness of MoMa, achieving superior performance with reduced computational cost.
title MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.23283