Bigger is not Always Better: Scaling Properties of Latent Diffusion Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Mei, Kangfu, Tu, Zhengzhong, Delbracio, Mauricio, Talebi, Hossein, Patel, Vishal M., Milanfar, Peyman
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910736715874304
author Mei, Kangfu
Tu, Zhengzhong
Delbracio, Mauricio
Talebi, Hossein
Patel, Vishal M.
Milanfar, Peyman
author_facet Mei, Kangfu
Tu, Zhengzhong
Delbracio, Mauricio
Talebi, Hossein
Patel, Vishal M.
Milanfar, Peyman
contents We study the scaling properties of latent diffusion models (LDMs) with an emphasis on their sampling efficiency. While improved network architecture and inference algorithms have shown to effectively boost sampling efficiency of diffusion models, the role of model size -- a critical determinant of sampling efficiency -- has not been thoroughly examined. Through empirical analysis of established text-to-image diffusion models, we conduct an in-depth investigation into how model size influences sampling efficiency across varying sampling steps. Our findings unveil a surprising trend: when operating under a given inference budget, smaller models frequently outperform their larger equivalents in generating high-quality results. Moreover, we extend our study to demonstrate the generalizability of the these findings by applying various diffusion samplers, exploring diverse downstream tasks, evaluating post-distilled models, as well as comparing performance relative to training compute. These findings open up new pathways for the development of LDM scaling strategies which can be employed to enhance generative capabilities within limited inference budgets.
format Preprint
id arxiv_https___arxiv_org_abs_2404_01367
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Bigger is not Always Better: Scaling Properties of Latent Diffusion Models
Mei, Kangfu
Tu, Zhengzhong
Delbracio, Mauricio
Talebi, Hossein
Patel, Vishal M.
Milanfar, Peyman
Computer Vision and Pattern Recognition
Machine Learning
We study the scaling properties of latent diffusion models (LDMs) with an emphasis on their sampling efficiency. While improved network architecture and inference algorithms have shown to effectively boost sampling efficiency of diffusion models, the role of model size -- a critical determinant of sampling efficiency -- has not been thoroughly examined. Through empirical analysis of established text-to-image diffusion models, we conduct an in-depth investigation into how model size influences sampling efficiency across varying sampling steps. Our findings unveil a surprising trend: when operating under a given inference budget, smaller models frequently outperform their larger equivalents in generating high-quality results. Moreover, we extend our study to demonstrate the generalizability of the these findings by applying various diffusion samplers, exploring diverse downstream tasks, evaluating post-distilled models, as well as comparing performance relative to training compute. These findings open up new pathways for the development of LDM scaling strategies which can be employed to enhance generative capabilities within limited inference budgets.
title Bigger is not Always Better: Scaling Properties of Latent Diffusion Models
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2404.01367