Generative Modeling of Weights: Generalization or Memorization?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zeng, Boya, Yin, Yida, Xu, Zhiqiu, Liu, Zhuang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916986459521024
author Zeng, Boya
Yin, Yida
Xu, Zhiqiu
Liu, Zhuang
author_facet Zeng, Boya
Yin, Yida
Xu, Zhiqiu
Liu, Zhuang
contents Generative models have recently been explored for synthesizing neural network weights. These approaches take neural network checkpoints as training data and aim to generate high-performing weights during inference. In this work, we examine four representative, well-known methods on their ability to generate novel model weights, i.e., weights that are different from the checkpoints seen during training. Contrary to claims in prior work, we find that these methods synthesize weights largely by memorization: they produce either replicas, or, at best, simple interpolations of the training checkpoints. Moreover, they fail to outperform simple baselines, such as adding noise to the weights or taking a simple weight ensemble, in obtaining different and simultaneously high-performing models. Our further analysis suggests that this memorization might result from limited data, overparameterized models, and the underuse of structural priors specific to weight data. These findings highlight the need for more careful design and rigorous evaluation of generative models when applied to new domains. Our code is available at https://github.com/boyazeng/weight_memorization.
format Preprint
id arxiv_https___arxiv_org_abs_2506_07998
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generative Modeling of Weights: Generalization or Memorization?
Zeng, Boya
Yin, Yida
Xu, Zhiqiu
Liu, Zhuang
Machine Learning
Computer Vision and Pattern Recognition
Generative models have recently been explored for synthesizing neural network weights. These approaches take neural network checkpoints as training data and aim to generate high-performing weights during inference. In this work, we examine four representative, well-known methods on their ability to generate novel model weights, i.e., weights that are different from the checkpoints seen during training. Contrary to claims in prior work, we find that these methods synthesize weights largely by memorization: they produce either replicas, or, at best, simple interpolations of the training checkpoints. Moreover, they fail to outperform simple baselines, such as adding noise to the weights or taking a simple weight ensemble, in obtaining different and simultaneously high-performing models. Our further analysis suggests that this memorization might result from limited data, overparameterized models, and the underuse of structural priors specific to weight data. These findings highlight the need for more careful design and rigorous evaluation of generative models when applied to new domains. Our code is available at https://github.com/boyazeng/weight_memorization.
title Generative Modeling of Weights: Generalization or Memorization?
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.07998