Information-Theoretic Framework for Understanding Modern Machine-Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feder, Meir, Urbanke, Ruediger, Fogel, Yaniv
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908623692627968
author Feder, Meir
Urbanke, Ruediger
Fogel, Yaniv
author_facet Feder, Meir
Urbanke, Ruediger
Fogel, Yaniv
contents We introduce an information-theoretic framework that views learning as universal prediction under log loss, characterized through regret bounds. Central to the framework is an effective notion of architecture-based model complexity, defined by the probability mass or volume of models in the vicinity of the data-generating process, or its projection on the model class. This volume is related to spectral properties of the expected Hessian or the Fisher Information Matrix, leading to tractable approximations. We argue that successful architectures possess a broad complexity range, enabling learning in highly over-parameterized model classes. The framework sheds light on the role of inductive biases, the effectiveness of stochastic gradient descent, and phenomena such as flat minima. It unifies online, batch, supervised, and generative settings, and applies across the stochastic-realizable and agnostic regimes. Moreover, it provides insights into the success of modern machine-learning architectures, such as deep neural networks and transformers, suggesting that their broad complexity range naturally arises from their layered structure. These insights open the door to the design of alternative architectures with potentially comparable or even superior performance.
format Preprint
id arxiv_https___arxiv_org_abs_2506_07661
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Information-Theoretic Framework for Understanding Modern Machine-Learning
Feder, Meir
Urbanke, Ruediger
Fogel, Yaniv
Machine Learning
Information Theory
We introduce an information-theoretic framework that views learning as universal prediction under log loss, characterized through regret bounds. Central to the framework is an effective notion of architecture-based model complexity, defined by the probability mass or volume of models in the vicinity of the data-generating process, or its projection on the model class. This volume is related to spectral properties of the expected Hessian or the Fisher Information Matrix, leading to tractable approximations. We argue that successful architectures possess a broad complexity range, enabling learning in highly over-parameterized model classes. The framework sheds light on the role of inductive biases, the effectiveness of stochastic gradient descent, and phenomena such as flat minima. It unifies online, batch, supervised, and generative settings, and applies across the stochastic-realizable and agnostic regimes. Moreover, it provides insights into the success of modern machine-learning architectures, such as deep neural networks and transformers, suggesting that their broad complexity range naturally arises from their layered structure. These insights open the door to the design of alternative architectures with potentially comparable or even superior performance.
title Information-Theoretic Framework for Understanding Modern Machine-Learning
topic Machine Learning
Information Theory
url https://arxiv.org/abs/2506.07661