Masked Autoencoders Are Effective Tokenizers for Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Hao, Han, Yujin, Chen, Fangyi, Li, Xiang, Wang, Yidong, Wang, Jindong, Wang, Ze, Liu, Zicheng, Zou, Difan, Raj, Bhiksha
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915312878747648
author Chen, Hao
Han, Yujin
Chen, Fangyi
Li, Xiang
Wang, Yidong
Wang, Jindong
Wang, Ze
Liu, Zicheng
Zou, Difan
Raj, Bhiksha
author_facet Chen, Hao
Han, Yujin
Chen, Fangyi
Li, Xiang
Wang, Yidong
Wang, Jindong
Wang, Ze
Liu, Zicheng
Zou, Difan
Raj, Bhiksha
contents Recent advances in latent diffusion models have demonstrated their effectiveness for high-resolution image synthesis. However, the properties of the latent space from tokenizer for better learning and generation of diffusion models remain under-explored. Theoretically and empirically, we find that improved generation quality is closely tied to the latent distributions with better structure, such as the ones with fewer Gaussian Mixture modes and more discriminative features. Motivated by these insights, we propose MAETok, an autoencoder (AE) leveraging mask modeling to learn semantically rich latent space while maintaining reconstruction fidelity. Extensive experiments validate our analysis, demonstrating that the variational form of autoencoders is not necessary, and a discriminative latent space from AE alone enables state-of-the-art performance on ImageNet generation using only 128 tokens. MAETok achieves significant practical improvements, enabling a gFID of 1.69 with 76x faster training and 31x higher inference throughput for 512x512 generation. Our findings show that the structure of the latent space, rather than variational constraints, is crucial for effective diffusion models. Code and trained models are released.
format Preprint
id arxiv_https___arxiv_org_abs_2502_03444
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Masked Autoencoders Are Effective Tokenizers for Diffusion Models
Chen, Hao
Han, Yujin
Chen, Fangyi
Li, Xiang
Wang, Yidong
Wang, Jindong
Wang, Ze
Liu, Zicheng
Zou, Difan
Raj, Bhiksha
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Recent advances in latent diffusion models have demonstrated their effectiveness for high-resolution image synthesis. However, the properties of the latent space from tokenizer for better learning and generation of diffusion models remain under-explored. Theoretically and empirically, we find that improved generation quality is closely tied to the latent distributions with better structure, such as the ones with fewer Gaussian Mixture modes and more discriminative features. Motivated by these insights, we propose MAETok, an autoencoder (AE) leveraging mask modeling to learn semantically rich latent space while maintaining reconstruction fidelity. Extensive experiments validate our analysis, demonstrating that the variational form of autoencoders is not necessary, and a discriminative latent space from AE alone enables state-of-the-art performance on ImageNet generation using only 128 tokens. MAETok achieves significant practical improvements, enabling a gFID of 1.69 with 76x faster training and 31x higher inference throughput for 512x512 generation. Our findings show that the structure of the latent space, rather than variational constraints, is crucial for effective diffusion models. Code and trained models are released.
title Masked Autoencoders Are Effective Tokenizers for Diffusion Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2502.03444