Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Choi, Hyesong, Kim, Daeun, Cha, Sungmin, Yi, Kwang Moo, Min, Dongbo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910764504186880
author Choi, Hyesong
Kim, Daeun
Cha, Sungmin
Yi, Kwang Moo
Min, Dongbo
author_facet Choi, Hyesong
Kim, Daeun
Cha, Sungmin
Yi, Kwang Moo
Min, Dongbo
contents In this work, we dive deep into the impact of additive noise in pre-training deep networks. While various methods have attempted to use additive noise inspired by the success of latent denoising diffusion models, when used in combination with masked image modeling, their gains have been marginal when it comes to recognition tasks. We thus investigate why this would be the case, in an attempt to find effective ways to combine the two ideas. Specifically, we find three critical conditions: corruption and restoration must be applied within the encoder, noise must be introduced in the feature space, and an explicit disentanglement between noised and masked tokens is necessary. By implementing these findings, we demonstrate improved pre-training performance for a wide range of recognition tasks, including those that require fine-grained, high-frequency information to solve.
format Preprint
id arxiv_https___arxiv_org_abs_2412_19104
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models
Choi, Hyesong
Kim, Daeun
Cha, Sungmin
Yi, Kwang Moo
Min, Dongbo
Computer Vision and Pattern Recognition
Machine Learning
In this work, we dive deep into the impact of additive noise in pre-training deep networks. While various methods have attempted to use additive noise inspired by the success of latent denoising diffusion models, when used in combination with masked image modeling, their gains have been marginal when it comes to recognition tasks. We thus investigate why this would be the case, in an attempt to find effective ways to combine the two ideas. Specifically, we find three critical conditions: corruption and restoration must be applied within the encoder, noise must be introduced in the feature space, and an explicit disentanglement between noised and masked tokens is necessary. By implementing these findings, we demonstrate improved pre-training performance for a wide range of recognition tasks, including those that require fine-grained, high-frequency information to solve.
title Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2412.19104