Saved in:
Bibliographic Details
Main Authors: Gao, Shanghua, Zhou, Pan, Cheng, Ming-Ming, Yan, Shuicheng
Format: Preprint
Published: 2023
Subjects:
Online Access:https://arxiv.org/abs/2303.14389
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914687125291008
author Gao, Shanghua
Zhou, Pan
Cheng, Ming-Ming
Yan, Shuicheng
author_facet Gao, Shanghua
Zhou, Pan
Cheng, Ming-Ming
Yan, Shuicheng
contents Despite its success in image synthesis, we observe that diffusion probabilistic models (DPMs) often lack contextual reasoning ability to learn the relations among object parts in an image, leading to a slow learning process. To solve this issue, we propose a Masked Diffusion Transformer (MDT) that introduces a mask latent modeling scheme to explicitly enhance the DPMs' ability to contextual relation learning among object semantic parts in an image. During training, MDT operates in the latent space to mask certain tokens. Then, an asymmetric diffusion transformer is designed to predict masked tokens from unmasked ones while maintaining the diffusion generation process. Our MDT can reconstruct the full information of an image from its incomplete contextual input, thus enabling it to learn the associated relations among image tokens. We further improve MDT with a more efficient macro network structure and training strategy, named MDTv2. Experimental results show that MDTv2 achieves superior image synthesis performance, e.g., a new SOTA FID score of 1.58 on the ImageNet dataset, and has more than 10x faster learning speed than the previous SOTA DiT. The source code is released at https://github.com/sail-sg/MDT.
format Preprint
id arxiv_https___arxiv_org_abs_2303_14389
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer
Gao, Shanghua
Zhou, Pan
Cheng, Ming-Ming
Yan, Shuicheng
Computer Vision and Pattern Recognition
Despite its success in image synthesis, we observe that diffusion probabilistic models (DPMs) often lack contextual reasoning ability to learn the relations among object parts in an image, leading to a slow learning process. To solve this issue, we propose a Masked Diffusion Transformer (MDT) that introduces a mask latent modeling scheme to explicitly enhance the DPMs' ability to contextual relation learning among object semantic parts in an image. During training, MDT operates in the latent space to mask certain tokens. Then, an asymmetric diffusion transformer is designed to predict masked tokens from unmasked ones while maintaining the diffusion generation process. Our MDT can reconstruct the full information of an image from its incomplete contextual input, thus enabling it to learn the associated relations among image tokens. We further improve MDT with a more efficient macro network structure and training strategy, named MDTv2. Experimental results show that MDTv2 achieves superior image synthesis performance, e.g., a new SOTA FID score of 1.58 on the ImageNet dataset, and has more than 10x faster learning speed than the previous SOTA DiT. The source code is released at https://github.com/sail-sg/MDT.
title MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2303.14389