Rethinking Patch Dependence for Masked Autoencoders

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fu, Letian, Lian, Long, Wang, Renhao, Shi, Baifeng, Wang, Xudong, Yala, Adam, Darrell, Trevor, Efros, Alexei A., Goldberg, Ken
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910907701919744
author Fu, Letian
Lian, Long
Wang, Renhao
Shi, Baifeng
Wang, Xudong
Yala, Adam
Darrell, Trevor
Efros, Alexei A.
Goldberg, Ken
author_facet Fu, Letian
Lian, Long
Wang, Renhao
Shi, Baifeng
Wang, Xudong
Yala, Adam
Darrell, Trevor
Efros, Alexei A.
Goldberg, Ken
contents In this work, we examine the impact of inter-patch dependencies in the decoder of masked autoencoders (MAE) on representation learning. We decompose the decoding mechanism for masked reconstruction into self-attention between mask tokens and cross-attention between masked and visible tokens. Our findings reveal that MAE reconstructs coherent images from visible patches not through interactions between patches in the decoder but by learning a global representation within the encoder. This discovery leads us to propose a simple visual pretraining framework: cross-attention masked autoencoders (CrossMAE). This framework employs only cross-attention in the decoder to independently read out reconstructions for a small subset of masked patches from encoder outputs. This approach achieves comparable or superior performance to traditional MAE across models ranging from ViT-S to ViT-H and significantly reduces computational requirements. By its design, CrossMAE challenges the necessity of interaction between mask tokens for effective masked pretraining. Code and models are publicly available: https://crossmae.github.io
format Preprint
id arxiv_https___arxiv_org_abs_2401_14391
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Rethinking Patch Dependence for Masked Autoencoders
Fu, Letian
Lian, Long
Wang, Renhao
Shi, Baifeng
Wang, Xudong
Yala, Adam
Darrell, Trevor
Efros, Alexei A.
Goldberg, Ken
Computer Vision and Pattern Recognition
In this work, we examine the impact of inter-patch dependencies in the decoder of masked autoencoders (MAE) on representation learning. We decompose the decoding mechanism for masked reconstruction into self-attention between mask tokens and cross-attention between masked and visible tokens. Our findings reveal that MAE reconstructs coherent images from visible patches not through interactions between patches in the decoder but by learning a global representation within the encoder. This discovery leads us to propose a simple visual pretraining framework: cross-attention masked autoencoders (CrossMAE). This framework employs only cross-attention in the decoder to independently read out reconstructions for a small subset of masked patches from encoder outputs. This approach achieves comparable or superior performance to traditional MAE across models ranging from ViT-S to ViT-H and significantly reduces computational requirements. By its design, CrossMAE challenges the necessity of interaction between mask tokens for effective masked pretraining. Code and models are publicly available: https://crossmae.github.io
title Rethinking Patch Dependence for Masked Autoencoders
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.14391