BIMM: Brain Inspired Masked Modeling for Video Representation Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wan, Zhifan, Zhang, Jie, Li, Changzhen, Shan, Shiguang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916254516772864
author Wan, Zhifan
Zhang, Jie
Li, Changzhen
Shan, Shiguang
author_facet Wan, Zhifan
Zhang, Jie
Li, Changzhen
Shan, Shiguang
contents The visual pathway of human brain includes two sub-pathways, ie, the ventral pathway and the dorsal pathway, which focus on object identification and dynamic information modeling, respectively. Both pathways comprise multi-layer structures, with each layer responsible for processing different aspects of visual information. Inspired by visual information processing mechanism of the human brain, we propose the Brain Inspired Masked Modeling (BIMM) framework, aiming to learn comprehensive representations from videos. Specifically, our approach consists of ventral and dorsal branches, which learn image and video representations, respectively. Both branches employ the Vision Transformer (ViT) as their backbone and are trained using masked modeling method. To achieve the goals of different visual cortices in the brain, we segment the encoder of each branch into three intermediate blocks and reconstruct progressive prediction targets with light weight decoders. Furthermore, drawing inspiration from the information-sharing mechanism in the visual pathways, we propose a partial parameter sharing strategy between the branches during training. Extensive experiments demonstrate that BIMM achieves superior performance compared to the state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2405_12757
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle BIMM: Brain Inspired Masked Modeling for Video Representation Learning
Wan, Zhifan
Zhang, Jie
Li, Changzhen
Shan, Shiguang
Computer Vision and Pattern Recognition
The visual pathway of human brain includes two sub-pathways, ie, the ventral pathway and the dorsal pathway, which focus on object identification and dynamic information modeling, respectively. Both pathways comprise multi-layer structures, with each layer responsible for processing different aspects of visual information. Inspired by visual information processing mechanism of the human brain, we propose the Brain Inspired Masked Modeling (BIMM) framework, aiming to learn comprehensive representations from videos. Specifically, our approach consists of ventral and dorsal branches, which learn image and video representations, respectively. Both branches employ the Vision Transformer (ViT) as their backbone and are trained using masked modeling method. To achieve the goals of different visual cortices in the brain, we segment the encoder of each branch into three intermediate blocks and reconstruct progressive prediction targets with light weight decoders. Furthermore, drawing inspiration from the information-sharing mechanism in the visual pathways, we propose a partial parameter sharing strategy between the branches during training. Extensive experiments demonstrate that BIMM achieves superior performance compared to the state-of-the-art methods.
title BIMM: Brain Inspired Masked Modeling for Video Representation Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.12757