Saved in:
Bibliographic Details
Main Authors: Xu, Yixian, Luo, Shengjie, Wang, Liwei, He, Di, Liu, Chang
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2506.13763
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917412392140800
author Xu, Yixian
Luo, Shengjie
Wang, Liwei
He, Di
Liu, Chang
author_facet Xu, Yixian
Luo, Shengjie
Wang, Liwei
He, Di
Liu, Chang
contents Diffusion models have achieved remarkable success in generative modeling. Despite more stable training, the loss of diffusion models is not indicative of absolute data-fitting quality, since its optimal value is typically not zero but unknown, leading to confusion between large optimal loss and insufficient model capacity. In this work, we advocate the need to estimate the optimal loss value for diagnosing and improving diffusion models. We first derive the optimal loss in closed form under a unified formulation of diffusion models, and develop effective estimators for it, including a stochastic variant scalable to large datasets with proper control of variance and bias. With this tool, we unlock the inherent metric for diagnosing the training quality of mainstream diffusion model variants, and develop a more performant training schedule based on the optimal loss. Moreover, using models with 120M to 1.5B parameters, we find that the power law is better demonstrated after subtracting the optimal loss from the actual training loss, suggesting a more principled setting for investigating the scaling law for diffusion models.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13763
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Diagnosing and Improving Diffusion Models by Estimating the Optimal Loss Value
Xu, Yixian
Luo, Shengjie
Wang, Liwei
He, Di
Liu, Chang
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Diffusion models have achieved remarkable success in generative modeling. Despite more stable training, the loss of diffusion models is not indicative of absolute data-fitting quality, since its optimal value is typically not zero but unknown, leading to confusion between large optimal loss and insufficient model capacity. In this work, we advocate the need to estimate the optimal loss value for diagnosing and improving diffusion models. We first derive the optimal loss in closed form under a unified formulation of diffusion models, and develop effective estimators for it, including a stochastic variant scalable to large datasets with proper control of variance and bias. With this tool, we unlock the inherent metric for diagnosing the training quality of mainstream diffusion model variants, and develop a more performant training schedule based on the optimal loss. Moreover, using models with 120M to 1.5B parameters, we find that the power law is better demonstrated after subtracting the optimal loss from the actual training loss, suggesting a more principled setting for investigating the scaling law for diffusion models.
title Diagnosing and Improving Diffusion Models by Estimating the Optimal Loss Value
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.13763