Saved in:
Bibliographic Details
Main Authors: Maleknia, Alex Alì, Sato, Yuzuru
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2604.02393
Tags: Add Tag
No Tags, Be the first to tag this record!
Table of Contents:
  • Vanishing gradients and overfitting are central problems in machine learning, yet are typically analyzed in asymptotic regimes that obscure their dynamical origins. Here we provide a dynamical description of learning in multi-layer perceptrons (MLPs) via a minimal model inspired by Fukumizu and Amari. We show that training dynamics traverse plateau and near-optimal regions, both organized by saddle structures, before converging to an overfitting regime. Under suitable conditions on the data, this regime collapses to a single attractor modulo symmetry. Furthermore, for finite noisy datasets, convergence to the theoretical optimum is impossible, and the dynamics necessarily settle into an overfitting solution.