NNiT: Width-Agnostic Neural Network Generation with Structurally Aligned Weight Spaces

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Jiwoo, Mehta, Swarajh, Hsu, Hao-Lun, Ryu, Hyunwoo, Liu, Yudong, Pajic, Miroslav
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915824277651456
author Kim, Jiwoo
Mehta, Swarajh
Hsu, Hao-Lun
Ryu, Hyunwoo
Liu, Yudong
Pajic, Miroslav
author_facet Kim, Jiwoo
Mehta, Swarajh
Hsu, Hao-Lun
Ryu, Hyunwoo
Liu, Yudong
Pajic, Miroslav
contents Generative modeling of neural network parameters is often tied to architectures because standard parameter representations rely on known weight-matrix dimensions. Generation is further complicated by permutation symmetries that allow networks to model similar input-output functions while having widely different, unaligned parameterizations. In this work, we introduce Neural Network Diffusion Transformers (NNiTs), which generate weights in a width-agnostic manner by tokenizing weight matrices into patches and modeling them as locally structured fields. We establish that Graph HyperNetworks (GHNs) with a convolutional neural network (CNN) decoder structurally align the weight space, creating the local correlation necessary for patch-based processing. Focusing on MLPs, where permutation symmetry is especially apparent, NNiT generates fully functional networks across a range of architectures. Our approach jointly models discrete architecture tokens and continuous weight patches within a single sequence model. On ManiSkill3 robotics tasks, NNiT achieves >85% success on architecture topologies unseen during training, while baseline approaches fail to generalize.
format Preprint
id arxiv_https___arxiv_org_abs_2603_00180
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle NNiT: Width-Agnostic Neural Network Generation with Structurally Aligned Weight Spaces
Kim, Jiwoo
Mehta, Swarajh
Hsu, Hao-Lun
Ryu, Hyunwoo
Liu, Yudong
Pajic, Miroslav
Machine Learning
Artificial Intelligence
Generative modeling of neural network parameters is often tied to architectures because standard parameter representations rely on known weight-matrix dimensions. Generation is further complicated by permutation symmetries that allow networks to model similar input-output functions while having widely different, unaligned parameterizations. In this work, we introduce Neural Network Diffusion Transformers (NNiTs), which generate weights in a width-agnostic manner by tokenizing weight matrices into patches and modeling them as locally structured fields. We establish that Graph HyperNetworks (GHNs) with a convolutional neural network (CNN) decoder structurally align the weight space, creating the local correlation necessary for patch-based processing. Focusing on MLPs, where permutation symmetry is especially apparent, NNiT generates fully functional networks across a range of architectures. Our approach jointly models discrete architecture tokens and continuous weight patches within a single sequence model. On ManiSkill3 robotics tasks, NNiT achieves >85% success on architecture topologies unseen during training, while baseline approaches fail to generalize.
title NNiT: Width-Agnostic Neural Network Generation with Structurally Aligned Weight Spaces
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2603.00180