Wasserstein Distances, Neuronal Entanglement, and Sparsity

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sawmya, Shashata, Kong, Linghao, Markov, Ilia, Alistarh, Dan, Shavit, Nir
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916631416930304
author Sawmya, Shashata
Kong, Linghao
Markov, Ilia
Alistarh, Dan
Shavit, Nir
author_facet Sawmya, Shashata
Kong, Linghao
Markov, Ilia
Alistarh, Dan
Shavit, Nir
contents Disentangling polysemantic neurons is at the core of many current approaches to interpretability of large language models. Here we attempt to study how disentanglement can be used to understand performance, particularly under weight sparsity, a leading post-training optimization technique. We suggest a novel measure for estimating neuronal entanglement: the Wasserstein distance of a neuron's output distribution to a Gaussian. Moreover, we show the existence of a small number of highly entangled "Wasserstein Neurons" in each linear layer of an LLM, characterized by their highly non-Gaussian output distributions, their role in mapping similar inputs to dissimilar outputs, and their significant impact on model accuracy. To study these phenomena, we propose a new experimental framework for disentangling polysemantic neurons. Our framework separates each layer's inputs to create a mixture of experts where each neuron's output is computed by a mixture of neurons of lower Wasserstein distance, each better at maintaining accuracy when sparsified without retraining. We provide strong evidence that this is because the mixture of sparse experts is effectively disentangling the input-output relationship of individual neurons, in particular the difficult Wasserstein neurons.
format Preprint
id arxiv_https___arxiv_org_abs_2405_15756
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Wasserstein Distances, Neuronal Entanglement, and Sparsity
Sawmya, Shashata
Kong, Linghao
Markov, Ilia
Alistarh, Dan
Shavit, Nir
Machine Learning
Artificial Intelligence
Disentangling polysemantic neurons is at the core of many current approaches to interpretability of large language models. Here we attempt to study how disentanglement can be used to understand performance, particularly under weight sparsity, a leading post-training optimization technique. We suggest a novel measure for estimating neuronal entanglement: the Wasserstein distance of a neuron's output distribution to a Gaussian. Moreover, we show the existence of a small number of highly entangled "Wasserstein Neurons" in each linear layer of an LLM, characterized by their highly non-Gaussian output distributions, their role in mapping similar inputs to dissimilar outputs, and their significant impact on model accuracy. To study these phenomena, we propose a new experimental framework for disentangling polysemantic neurons. Our framework separates each layer's inputs to create a mixture of experts where each neuron's output is computed by a mixture of neurons of lower Wasserstein distance, each better at maintaining accuracy when sparsified without retraining. We provide strong evidence that this is because the mixture of sparse experts is effectively disentangling the input-output relationship of individual neurons, in particular the difficult Wasserstein neurons.
title Wasserstein Distances, Neuronal Entanglement, and Sparsity
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2405.15756