The Butterfly Effect: Neural Network Training Trajectories Are Highly Sensitive to Initial Conditions

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kwok, Devin, Altıntaş, Gül Sena, Raffel, Colin, Rolnick, David
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909650056642560
author Kwok, Devin
Altıntaş, Gül Sena
Raffel, Colin
Rolnick, David
author_facet Kwok, Devin
Altıntaş, Gül Sena
Raffel, Colin
Rolnick, David
contents Neural network training is inherently sensitive to initialization and the randomness induced by stochastic gradient descent. However, it is unclear to what extent such effects lead to meaningfully different networks, either in terms of the models' weights or the underlying functions that were learned. In this work, we show that during the initial "chaotic" phase of training, even extremely small perturbations reliably causes otherwise identical training trajectories to diverge-an effect that diminishes rapidly over training time. We quantify this divergence through (i) $L^2$ distance between parameters, (ii) the loss barrier when interpolating between networks, (iii) $L^2$ and barrier between parameters after permutation alignment, and (iv) representational similarity between intermediate activations; revealing how perturbations across different hyperparameter or fine-tuning settings drive training trajectories toward distinct loss minima. Our findings provide insights into neural network training stability, with practical implications for fine-tuning, model merging, and diversity of model ensembles.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13234
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Butterfly Effect: Neural Network Training Trajectories Are Highly Sensitive to Initial Conditions
Kwok, Devin
Altıntaş, Gül Sena
Raffel, Colin
Rolnick, David
Machine Learning
Neural network training is inherently sensitive to initialization and the randomness induced by stochastic gradient descent. However, it is unclear to what extent such effects lead to meaningfully different networks, either in terms of the models' weights or the underlying functions that were learned. In this work, we show that during the initial "chaotic" phase of training, even extremely small perturbations reliably causes otherwise identical training trajectories to diverge-an effect that diminishes rapidly over training time. We quantify this divergence through (i) $L^2$ distance between parameters, (ii) the loss barrier when interpolating between networks, (iii) $L^2$ and barrier between parameters after permutation alignment, and (iv) representational similarity between intermediate activations; revealing how perturbations across different hyperparameter or fine-tuning settings drive training trajectories toward distinct loss minima. Our findings provide insights into neural network training stability, with practical implications for fine-tuning, model merging, and diversity of model ensembles.
title The Butterfly Effect: Neural Network Training Trajectories Are Highly Sensitive to Initial Conditions
topic Machine Learning
url https://arxiv.org/abs/2506.13234