Uncovering a Universal Abstract Algorithm for Modular Addition in Neural Networks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: McCracken, Gavin, Moisescu-Pareja, Gabriela, Letourneau, Vincent, Precup, Doina, Love, Jonathan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913857438482432
author McCracken, Gavin
Moisescu-Pareja, Gabriela
Letourneau, Vincent
Precup, Doina
Love, Jonathan
author_facet McCracken, Gavin
Moisescu-Pareja, Gabriela
Letourneau, Vincent
Precup, Doina
Love, Jonathan
contents We propose a testable universality hypothesis, asserting that seemingly disparate neural network solutions observed in the simple task of modular addition are unified under a common abstract algorithm. While prior work interpreted variations in neuron-level representations as evidence for distinct algorithms, we demonstrate - through multi-level analyses spanning neurons, neuron clusters, and entire networks - that multilayer perceptrons and transformers universally implement the abstract algorithm we call the approximate Chinese Remainder Theorem. Crucially, we introduce approximate cosets and show that neurons activate exclusively on them. Furthermore, our theory works for deep neural networks (DNNs). It predicts that universally learned solutions in DNNs with trainable embeddings or more than one hidden layer require only O(log n) features, a result we empirically confirm. This work thus provides the first theory-backed interpretation of multilayer networks solving modular addition. It advances generalizable interpretability and opens a testable universality hypothesis for group multiplication beyond modular addition.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18266
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Uncovering a Universal Abstract Algorithm for Modular Addition in Neural Networks
McCracken, Gavin
Moisescu-Pareja, Gabriela
Letourneau, Vincent
Precup, Doina
Love, Jonathan
Machine Learning
Artificial Intelligence
We propose a testable universality hypothesis, asserting that seemingly disparate neural network solutions observed in the simple task of modular addition are unified under a common abstract algorithm. While prior work interpreted variations in neuron-level representations as evidence for distinct algorithms, we demonstrate - through multi-level analyses spanning neurons, neuron clusters, and entire networks - that multilayer perceptrons and transformers universally implement the abstract algorithm we call the approximate Chinese Remainder Theorem. Crucially, we introduce approximate cosets and show that neurons activate exclusively on them. Furthermore, our theory works for deep neural networks (DNNs). It predicts that universally learned solutions in DNNs with trainable embeddings or more than one hidden layer require only O(log n) features, a result we empirically confirm. This work thus provides the first theory-backed interpretation of multilayer networks solving modular addition. It advances generalizable interpretability and opens a testable universality hypothesis for group multiplication beyond modular addition.
title Uncovering a Universal Abstract Algorithm for Modular Addition in Neural Networks
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.18266