Deconstructing Denoising Diffusion Models for Self-Supervised Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Xinlei, Liu, Zhuang, Xie, Saining, He, Kaiming
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929224416231424
author Chen, Xinlei
Liu, Zhuang
Xie, Saining
He, Kaiming
author_facet Chen, Xinlei
Liu, Zhuang
Xie, Saining
He, Kaiming
contents In this study, we examine the representation learning abilities of Denoising Diffusion Models (DDM) that were originally purposed for image generation. Our philosophy is to deconstruct a DDM, gradually transforming it into a classical Denoising Autoencoder (DAE). This deconstructive procedure allows us to explore how various components of modern DDMs influence self-supervised representation learning. We observe that only a very few modern components are critical for learning good representations, while many others are nonessential. Our study ultimately arrives at an approach that is highly simplified and to a large extent resembles a classical DAE. We hope our study will rekindle interest in a family of classical methods within the realm of modern self-supervised learning.
format Preprint
id arxiv_https___arxiv_org_abs_2401_14404
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Deconstructing Denoising Diffusion Models for Self-Supervised Learning
Chen, Xinlei
Liu, Zhuang
Xie, Saining
He, Kaiming
Computer Vision and Pattern Recognition
Machine Learning
In this study, we examine the representation learning abilities of Denoising Diffusion Models (DDM) that were originally purposed for image generation. Our philosophy is to deconstruct a DDM, gradually transforming it into a classical Denoising Autoencoder (DAE). This deconstructive procedure allows us to explore how various components of modern DDMs influence self-supervised representation learning. We observe that only a very few modern components are critical for learning good representations, while many others are nonessential. Our study ultimately arrives at an approach that is highly simplified and to a large extent resembles a classical DAE. We hope our study will rekindle interest in a family of classical methods within the realm of modern self-supervised learning.
title Deconstructing Denoising Diffusion Models for Self-Supervised Learning
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2401.14404