USP: Unified Self-Supervised Pretraining for Image Generation and Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chu, Xiangxiang, Li, Renda, Wang, Yong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908459949096960
author Chu, Xiangxiang
Li, Renda
Wang, Yong
author_facet Chu, Xiangxiang
Li, Renda
Wang, Yong
contents Recent studies have highlighted the interplay between diffusion models and representation learning. Intermediate representations from diffusion models can be leveraged for downstream visual tasks, while self-supervised vision models can enhance the convergence and generation quality of diffusion models. However, transferring pretrained weights from vision models to diffusion models is challenging due to input mismatches and the use of latent spaces. To address these challenges, we propose Unified Self-supervised Pretraining (USP), a framework that initializes diffusion models via masked latent modeling in a Variational Autoencoder (VAE) latent space. USP achieves comparable performance in understanding tasks while significantly improving the convergence speed and generation quality of diffusion models. Our code will be publicly available at https://github.com/AMAP-ML/USP.
format Preprint
id arxiv_https___arxiv_org_abs_2503_06132
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle USP: Unified Self-Supervised Pretraining for Image Generation and Understanding
Chu, Xiangxiang
Li, Renda
Wang, Yong
Computer Vision and Pattern Recognition
Recent studies have highlighted the interplay between diffusion models and representation learning. Intermediate representations from diffusion models can be leveraged for downstream visual tasks, while self-supervised vision models can enhance the convergence and generation quality of diffusion models. However, transferring pretrained weights from vision models to diffusion models is challenging due to input mismatches and the use of latent spaces. To address these challenges, we propose Unified Self-supervised Pretraining (USP), a framework that initializes diffusion models via masked latent modeling in a Variational Autoencoder (VAE) latent space. USP achieves comparable performance in understanding tasks while significantly improving the convergence speed and generation quality of diffusion models. Our code will be publicly available at https://github.com/AMAP-ML/USP.
title USP: Unified Self-Supervised Pretraining for Image Generation and Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.06132