Not All Parameters Matter: Masking Diffusion Models for Enhancing Generation Ability

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Lei, Li, Senmao, Yang, Fei, Wang, Jianye, Zhang, Ziheng, Liu, Yuhan, Wang, Yaxing, Yang, Jian
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916721970905088
author Wang, Lei
Li, Senmao
Yang, Fei
Wang, Jianye
Zhang, Ziheng
Liu, Yuhan
Wang, Yaxing
Yang, Jian
author_facet Wang, Lei
Li, Senmao
Yang, Fei
Wang, Jianye
Zhang, Ziheng
Liu, Yuhan
Wang, Yaxing
Yang, Jian
contents The diffusion models, in early stages focus on constructing basic image structures, while the refined details, including local features and textures, are generated in later stages. Thus the same network layers are forced to learn both structural and textural information simultaneously, significantly differing from the traditional deep learning architectures (e.g., ResNet or GANs) which captures or generates the image semantic information at different layers. This difference inspires us to explore the time-wise diffusion models. We initially investigate the key contributions of the U-Net parameters to the denoising process and identify that properly zeroing out certain parameters (including large parameters) contributes to denoising, substantially improving the generation quality on the fly. Capitalizing on this discovery, we propose a simple yet effective method-termed ``MaskUNet''- that enhances generation quality with negligible parameter numbers. Our method fully leverages timestep- and sample-dependent effective U-Net parameters. To optimize MaskUNet, we offer two fine-tuning strategies: a training-based approach and a training-free approach, including tailored networks and optimization functions. In zero-shot inference on the COCO dataset, MaskUNet achieves the best FID score and further demonstrates its effectiveness in downstream task evaluations. Project page: https://gudaochangsheng.github.io/MaskUnet-Page/
format Preprint
id arxiv_https___arxiv_org_abs_2505_03097
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Not All Parameters Matter: Masking Diffusion Models for Enhancing Generation Ability
Wang, Lei
Li, Senmao
Yang, Fei
Wang, Jianye
Zhang, Ziheng
Liu, Yuhan
Wang, Yaxing
Yang, Jian
Computer Vision and Pattern Recognition
The diffusion models, in early stages focus on constructing basic image structures, while the refined details, including local features and textures, are generated in later stages. Thus the same network layers are forced to learn both structural and textural information simultaneously, significantly differing from the traditional deep learning architectures (e.g., ResNet or GANs) which captures or generates the image semantic information at different layers. This difference inspires us to explore the time-wise diffusion models. We initially investigate the key contributions of the U-Net parameters to the denoising process and identify that properly zeroing out certain parameters (including large parameters) contributes to denoising, substantially improving the generation quality on the fly. Capitalizing on this discovery, we propose a simple yet effective method-termed ``MaskUNet''- that enhances generation quality with negligible parameter numbers. Our method fully leverages timestep- and sample-dependent effective U-Net parameters. To optimize MaskUNet, we offer two fine-tuning strategies: a training-based approach and a training-free approach, including tailored networks and optimization functions. In zero-shot inference on the COCO dataset, MaskUNet achieves the best FID score and further demonstrates its effectiveness in downstream task evaluations. Project page: https://gudaochangsheng.github.io/MaskUnet-Page/
title Not All Parameters Matter: Masking Diffusion Models for Enhancing Generation Ability
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.03097