Masked Diffusion as Self-supervised Representation Learner

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Pan, Zixuan, Chen, Jianxu, Shi, Yiyu
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916203997429760
author Pan, Zixuan
Chen, Jianxu
Shi, Yiyu
author_facet Pan, Zixuan
Chen, Jianxu
Shi, Yiyu
contents Denoising diffusion probabilistic models have recently demonstrated state-of-the-art generative performance and have been used as strong pixel-level representation learners. This paper decomposes the interrelation between the generative capability and representation learning ability inherent in diffusion models. We present the masked diffusion model (MDM), a scalable self-supervised representation learner for semantic segmentation, substituting the conventional additive Gaussian noise of traditional diffusion with a masking mechanism. Our proposed approach convincingly surpasses prior benchmarks, demonstrating remarkable advancements in both medical and natural image semantic segmentation tasks, particularly in few-shot scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2308_05695
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Masked Diffusion as Self-supervised Representation Learner
Pan, Zixuan
Chen, Jianxu
Shi, Yiyu
Computer Vision and Pattern Recognition
Denoising diffusion probabilistic models have recently demonstrated state-of-the-art generative performance and have been used as strong pixel-level representation learners. This paper decomposes the interrelation between the generative capability and representation learning ability inherent in diffusion models. We present the masked diffusion model (MDM), a scalable self-supervised representation learner for semantic segmentation, substituting the conventional additive Gaussian noise of traditional diffusion with a masking mechanism. Our proposed approach convincingly surpasses prior benchmarks, demonstrating remarkable advancements in both medical and natural image semantic segmentation tasks, particularly in few-shot scenarios.
title Masked Diffusion as Self-supervised Representation Learner
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2308.05695