Binarized Neural Networks Converge Toward Algorithmic Simplicity: Empirical Support for the Learning-as-Compression Hypothesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sakabe, Eduardo Y., Abrahão, Felipe S., Simões, Alexandre, Colombini, Esther, Costa, Paula, Gudwin, Ricardo, Zenil, Hector
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909795614720000
author Sakabe, Eduardo Y.
Abrahão, Felipe S.
Simões, Alexandre
Colombini, Esther
Costa, Paula
Gudwin, Ricardo
Zenil, Hector
author_facet Sakabe, Eduardo Y.
Abrahão, Felipe S.
Simões, Alexandre
Colombini, Esther
Costa, Paula
Gudwin, Ricardo
Zenil, Hector
contents Understanding and controlling the informational complexity of neural networks is a central challenge in machine learning, with implications for generalization, optimization, and model capacity. While most approaches rely on entropy-based loss functions and statistical metrics, these measures often fail to capture deeper, causally relevant algorithmic regularities embedded in network structure. We propose a shift toward algorithmic information theory, using Binarized Neural Networks (BNNs) as a first proxy. Grounded in algorithmic probability (AP) and the universal distribution it defines, our approach characterizes learning dynamics through a formal, causally grounded lens. We apply the Block Decomposition Method (BDM) -- a scalable approximation of algorithmic complexity based on AP -- and demonstrate that it more closely tracks structural changes during training than entropy, consistently exhibiting stronger correlations with training loss across varying model sizes and randomized training runs. These results support the view of training as a process of algorithmic compression, where learning corresponds to the progressive internalization of structured regularities. In doing so, our work offers a principled estimate of learning progression and suggests a framework for complexity-aware learning and regularization, grounded in first principles from information theory, complexity, and computability.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20646
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Binarized Neural Networks Converge Toward Algorithmic Simplicity: Empirical Support for the Learning-as-Compression Hypothesis
Sakabe, Eduardo Y.
Abrahão, Felipe S.
Simões, Alexandre
Colombini, Esther
Costa, Paula
Gudwin, Ricardo
Zenil, Hector
Machine Learning
Artificial Intelligence
Information Theory
68T07, 68Q30, 68Q32
I.2.6; F.1.1; F.1.3
Understanding and controlling the informational complexity of neural networks is a central challenge in machine learning, with implications for generalization, optimization, and model capacity. While most approaches rely on entropy-based loss functions and statistical metrics, these measures often fail to capture deeper, causally relevant algorithmic regularities embedded in network structure. We propose a shift toward algorithmic information theory, using Binarized Neural Networks (BNNs) as a first proxy. Grounded in algorithmic probability (AP) and the universal distribution it defines, our approach characterizes learning dynamics through a formal, causally grounded lens. We apply the Block Decomposition Method (BDM) -- a scalable approximation of algorithmic complexity based on AP -- and demonstrate that it more closely tracks structural changes during training than entropy, consistently exhibiting stronger correlations with training loss across varying model sizes and randomized training runs. These results support the view of training as a process of algorithmic compression, where learning corresponds to the progressive internalization of structured regularities. In doing so, our work offers a principled estimate of learning progression and suggests a framework for complexity-aware learning and regularization, grounded in first principles from information theory, complexity, and computability.
title Binarized Neural Networks Converge Toward Algorithmic Simplicity: Empirical Support for the Learning-as-Compression Hypothesis
topic Machine Learning
Artificial Intelligence
Information Theory
68T07, 68Q30, 68Q32
I.2.6; F.1.1; F.1.3
url https://arxiv.org/abs/2505.20646