PiLaMIM: Toward Richer Visual Representations by Integrating Pixel and Latent Masked Image Modeling

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lee, Junmyeong, Hwang, Eui Jun, Cho, Sukmin, Park, Jong C.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915092414595072
author Lee, Junmyeong
Hwang, Eui Jun
Cho, Sukmin
Park, Jong C.
author_facet Lee, Junmyeong
Hwang, Eui Jun
Cho, Sukmin
Park, Jong C.
contents In Masked Image Modeling (MIM), two primary methods exist: Pixel MIM and Latent MIM, each utilizing different reconstruction targets, raw pixels and latent representations, respectively. Pixel MIM tends to capture low-level visual details such as color and texture, while Latent MIM focuses on high-level semantics of an object. However, these distinct strengths of each method can lead to suboptimal performance in tasks that rely on a particular level of visual features. To address this limitation, we propose PiLaMIM, a unified framework that combines Pixel MIM and Latent MIM to integrate their complementary strengths. Our method uses a single encoder along with two distinct decoders: one for predicting pixel values and another for latent representations, ensuring the capture of both high-level and low-level visual features. We further integrate the CLS token into the reconstruction process to aggregate global context, enabling the model to capture more semantic information. Extensive experiments demonstrate that PiLaMIM outperforms key baselines such as MAE, I-JEPA and BootMAE in most cases, proving its effectiveness in extracting richer visual representations.
format Preprint
id arxiv_https___arxiv_org_abs_2501_03005
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PiLaMIM: Toward Richer Visual Representations by Integrating Pixel and Latent Masked Image Modeling
Lee, Junmyeong
Hwang, Eui Jun
Cho, Sukmin
Park, Jong C.
Computer Vision and Pattern Recognition
In Masked Image Modeling (MIM), two primary methods exist: Pixel MIM and Latent MIM, each utilizing different reconstruction targets, raw pixels and latent representations, respectively. Pixel MIM tends to capture low-level visual details such as color and texture, while Latent MIM focuses on high-level semantics of an object. However, these distinct strengths of each method can lead to suboptimal performance in tasks that rely on a particular level of visual features. To address this limitation, we propose PiLaMIM, a unified framework that combines Pixel MIM and Latent MIM to integrate their complementary strengths. Our method uses a single encoder along with two distinct decoders: one for predicting pixel values and another for latent representations, ensuring the capture of both high-level and low-level visual features. We further integrate the CLS token into the reconstruction process to aggregate global context, enabling the model to capture more semantic information. Extensive experiments demonstrate that PiLaMIM outperforms key baselines such as MAE, I-JEPA and BootMAE in most cases, proving its effectiveness in extracting richer visual representations.
title PiLaMIM: Toward Richer Visual Representations by Integrating Pixel and Latent Masked Image Modeling
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.03005