Studying Cross-cluster Modularity in Neural Networks

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Golechha, Satvik, Chaudhary, Maheep, Velja, Joan, Abate, Alessandro, Schoots, Nandi
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911076711399424
author Golechha, Satvik
Chaudhary, Maheep
Velja, Joan
Abate, Alessandro
Schoots, Nandi
author_facet Golechha, Satvik
Chaudhary, Maheep
Velja, Joan
Abate, Alessandro
Schoots, Nandi
contents An approach to improve neural network interpretability is via clusterability, i.e., splitting a model into disjoint clusters that can be studied independently. We define a measure for clusterability and show that pre-trained models form highly enmeshed clusters via spectral graph clustering. We thus train models to be more modular using a "clusterability loss" function that encourages the formation of non-interacting clusters. We then investigate the emerging properties of these highly clustered models. We find our trained clustered models do not exhibit more task specialization, but do form smaller circuits. We investigate CNNs trained on MNIST and CIFAR, small transformers trained on modular addition, and GPT-2 and Pythia on the Wiki dataset, and Gemma on a Chemistry dataset. This investigation shows what to expect from clustered models.
format Preprint
id arxiv_https___arxiv_org_abs_2502_02470
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Studying Cross-cluster Modularity in Neural Networks
Golechha, Satvik
Chaudhary, Maheep
Velja, Joan
Abate, Alessandro
Schoots, Nandi
Machine Learning
Artificial Intelligence
An approach to improve neural network interpretability is via clusterability, i.e., splitting a model into disjoint clusters that can be studied independently. We define a measure for clusterability and show that pre-trained models form highly enmeshed clusters via spectral graph clustering. We thus train models to be more modular using a "clusterability loss" function that encourages the formation of non-interacting clusters. We then investigate the emerging properties of these highly clustered models. We find our trained clustered models do not exhibit more task specialization, but do form smaller circuits. We investigate CNNs trained on MNIST and CIFAR, small transformers trained on modular addition, and GPT-2 and Pythia on the Wiki dataset, and Gemma on a Chemistry dataset. This investigation shows what to expect from clustered models.
title Studying Cross-cluster Modularity in Neural Networks
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2502.02470