Generalization of Diffusion Models Arises with a Balanced Representation Space

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zekai, Li, Xiao, Li, Xiang, Shi, Lianghe, Wu, Meng, Tao, Molei, Qu, Qing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911439062564864
author Zhang, Zekai
Li, Xiao
Li, Xiang
Shi, Lianghe
Wu, Meng
Tao, Molei
Qu, Qing
author_facet Zhang, Zekai
Li, Xiao
Li, Xiang
Shi, Lianghe
Wu, Meng
Tao, Molei
Qu, Qing
contents Diffusion models excel at generating high-quality, diverse samples, yet they risk memorizing training data when overfit to the training objective. We analyze the distinctions between memorization and generalization in diffusion models through the lens of representation learning. By investigating a two-layer ReLU denoising autoencoder (DAE), we prove that (i) memorization corresponds to the model storing raw training samples in the learned weights for encoding and decoding, yielding localized spiky representations, whereas (ii) generalization arises when the model captures local data statistics, producing balanced representations. Furthermore, we validate these theoretical findings on real-world unconditional and text-to-image diffusion models, demonstrating that the same representation structures emerge in deep generative models with significant practical implications. Building on these insights, we propose a representation-based method for detecting memorization and a training-free editing technique that allows precise control via representation steering. Together, our results highlight that learning good representations is central to novel and meaningful generative modeling.
format Preprint
id arxiv_https___arxiv_org_abs_2512_20963
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generalization of Diffusion Models Arises with a Balanced Representation Space
Zhang, Zekai
Li, Xiao
Li, Xiang
Shi, Lianghe
Wu, Meng
Tao, Molei
Qu, Qing
Machine Learning
Computer Vision and Pattern Recognition
Diffusion models excel at generating high-quality, diverse samples, yet they risk memorizing training data when overfit to the training objective. We analyze the distinctions between memorization and generalization in diffusion models through the lens of representation learning. By investigating a two-layer ReLU denoising autoencoder (DAE), we prove that (i) memorization corresponds to the model storing raw training samples in the learned weights for encoding and decoding, yielding localized spiky representations, whereas (ii) generalization arises when the model captures local data statistics, producing balanced representations. Furthermore, we validate these theoretical findings on real-world unconditional and text-to-image diffusion models, demonstrating that the same representation structures emerge in deep generative models with significant practical implications. Building on these insights, we propose a representation-based method for detecting memorization and a training-free editing technique that allows precise control via representation steering. Together, our results highlight that learning good representations is central to novel and meaningful generative modeling.
title Generalization of Diffusion Models Arises with a Balanced Representation Space
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.20963