From Neurons to Neutrons: A Case Study in Interpretability

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kitouni, Ouail, Nolte, Niklas, Pérez-Díaz, Víctor Samuel, Trifinopoulos, Sokratis, Williams, Mike
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911889350459392
author Kitouni, Ouail
Nolte, Niklas
Pérez-Díaz, Víctor Samuel
Trifinopoulos, Sokratis
Williams, Mike
author_facet Kitouni, Ouail
Nolte, Niklas
Pérez-Díaz, Víctor Samuel
Trifinopoulos, Sokratis
Williams, Mike
contents Mechanistic Interpretability (MI) promises a path toward fully understanding how neural networks make their predictions. Prior work demonstrates that even when trained to perform simple arithmetic, models can implement a variety of algorithms (sometimes concurrently) depending on initialization and hyperparameters. Does this mean neuron-level interpretability techniques have limited applicability? We argue that high-dimensional neural networks can learn low-dimensional representations of their training data that are useful beyond simply making good predictions. Such representations can be understood through the mechanistic interpretability lens and provide insights that are surprisingly faithful to human-derived domain knowledge. This indicates that such approaches to interpretability can be useful for deriving a new understanding of a problem from models trained to solve it. As a case study, we extract nuclear physics concepts by studying models trained to reproduce nuclear data.
format Preprint
id arxiv_https___arxiv_org_abs_2405_17425
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle From Neurons to Neutrons: A Case Study in Interpretability
Kitouni, Ouail
Nolte, Niklas
Pérez-Díaz, Víctor Samuel
Trifinopoulos, Sokratis
Williams, Mike
Machine Learning
Nuclear Theory
Mechanistic Interpretability (MI) promises a path toward fully understanding how neural networks make their predictions. Prior work demonstrates that even when trained to perform simple arithmetic, models can implement a variety of algorithms (sometimes concurrently) depending on initialization and hyperparameters. Does this mean neuron-level interpretability techniques have limited applicability? We argue that high-dimensional neural networks can learn low-dimensional representations of their training data that are useful beyond simply making good predictions. Such representations can be understood through the mechanistic interpretability lens and provide insights that are surprisingly faithful to human-derived domain knowledge. This indicates that such approaches to interpretability can be useful for deriving a new understanding of a problem from models trained to solve it. As a case study, we extract nuclear physics concepts by studying models trained to reproduce nuclear data.
title From Neurons to Neutrons: A Case Study in Interpretability
topic Machine Learning
Nuclear Theory
url https://arxiv.org/abs/2405.17425