R-MAE: Regions Meet Masked Autoencoders

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Duy-Kien, Aggarwal, Vaibhav, Li, Yanghao, Oswald, Martin R., Kirillov, Alexander, Snoek, Cees G. M., Chen, Xinlei
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910287591899136
author Nguyen, Duy-Kien
Aggarwal, Vaibhav
Li, Yanghao
Oswald, Martin R.
Kirillov, Alexander
Snoek, Cees G. M.
Chen, Xinlei
author_facet Nguyen, Duy-Kien
Aggarwal, Vaibhav
Li, Yanghao
Oswald, Martin R.
Kirillov, Alexander
Snoek, Cees G. M.
Chen, Xinlei
contents In this work, we explore regions as a potential visual analogue of words for self-supervised image representation learning. Inspired by Masked Autoencoding (MAE), a generative pre-training baseline, we propose masked region autoencoding to learn from groups of pixels or regions. Specifically, we design an architecture which efficiently addresses the one-to-many mapping between images and regions, while being highly effective especially with high-quality regions. When integrated with MAE, our approach (R-MAE) demonstrates consistent improvements across various pre-training datasets and downstream detection and segmentation benchmarks, with negligible computational overheads. Beyond the quantitative evaluation, our analysis indicates the models pre-trained with masked region autoencoding unlock the potential for interactive segmentation. The code is provided at https://github.com/facebookresearch/r-mae.
format Preprint
id arxiv_https___arxiv_org_abs_2306_05411
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle R-MAE: Regions Meet Masked Autoencoders
Nguyen, Duy-Kien
Aggarwal, Vaibhav
Li, Yanghao
Oswald, Martin R.
Kirillov, Alexander
Snoek, Cees G. M.
Chen, Xinlei
Computer Vision and Pattern Recognition
In this work, we explore regions as a potential visual analogue of words for self-supervised image representation learning. Inspired by Masked Autoencoding (MAE), a generative pre-training baseline, we propose masked region autoencoding to learn from groups of pixels or regions. Specifically, we design an architecture which efficiently addresses the one-to-many mapping between images and regions, while being highly effective especially with high-quality regions. When integrated with MAE, our approach (R-MAE) demonstrates consistent improvements across various pre-training datasets and downstream detection and segmentation benchmarks, with negligible computational overheads. Beyond the quantitative evaluation, our analysis indicates the models pre-trained with masked region autoencoding unlock the potential for interactive segmentation. The code is provided at https://github.com/facebookresearch/r-mae.
title R-MAE: Regions Meet Masked Autoencoders
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2306.05411