The Butterfly Effect: Neural Network Training Trajectories Are Highly Sensitive to Initial Conditions
Fuente:
arXiv
Saved in:
| Main Authors: | Kwok, Devin, Altıntaş, Gül Sena, Raffel, Colin, Rolnick, David |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TokSuite: Measuring the Impact of Tokenizer Choice on Language Model Behavior
by: Altıntaş, Gül Sena, et al.
Published: (2025)
by: Altıntaş, Gül Sena, et al.
Published: (2025)
Dataset Difficulty and the Role of Inductive Bias
by: Kwok, Devin, et al.
Published: (2024)
by: Kwok, Devin, et al.
Published: (2024)
Position: The Most Expensive Part of an LLM should be its Training Data
by: Kandpal, Nikhil, et al.
Published: (2025)
by: Kandpal, Nikhil, et al.
Published: (2025)
Uncovering Model Processing Strategies with Non-Negative Per-Example Fisher Factorization
by: Matena, Michael, et al.
Published: (2023)
by: Matena, Michael, et al.
Published: (2023)
Simultaneous linear connectivity of neural networks modulo permutation
by: Sharma, Ekansh, et al.
Published: (2024)
by: Sharma, Ekansh, et al.
Published: (2024)
FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale
by: Patel, Ajay, et al.
Published: (2026)
by: Patel, Ajay, et al.
Published: (2026)
Towards Climate Variable Prediction with Conditioned Spatio-Temporal Normalizing Flows
by: Winkler, Christina, et al.
Published: (2023)
by: Winkler, Christina, et al.
Published: (2023)
Efficiently Estimating Data Efficiency for Language Model Fine-tuning
by: Je, Gyung Hyun, et al.
Published: (2025)
by: Je, Gyung Hyun, et al.
Published: (2025)
Enhancing Training Data Attribution with Representational Optimization
by: Sun, Weiwei, et al.
Published: (2025)
by: Sun, Weiwei, et al.
Published: (2025)
Effects of Initialization Biases on Deep Neural Network Training Dynamics
by: Pellegrino, Nicholas, et al.
Published: (2025)
by: Pellegrino, Nicholas, et al.
Published: (2025)
AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution
by: Liu, Fengyuan, et al.
Published: (2024)
by: Liu, Fengyuan, et al.
Published: (2024)
Soft Merging of Experts with Adaptive Routing
by: Muqeeth, Mohammed, et al.
Published: (2023)
by: Muqeeth, Mohammed, et al.
Published: (2023)
Merging by Matching Models in Task Parameter Subspaces
by: Tam, Derek, et al.
Published: (2023)
by: Tam, Derek, et al.
Published: (2023)
Learning to Route Among Specialized Experts for Zero-Shot Generalization
by: Muqeeth, Mohammed, et al.
Published: (2024)
by: Muqeeth, Mohammed, et al.
Published: (2024)
Depth-Aware Initialization for Stable and Efficient Neural Network Training
by: Pandey, Vijay
Published: (2025)
by: Pandey, Vijay
Published: (2025)
DataDreamer: A Tool for Synthetic Data Generation and Reproducible LLM Workflows
by: Patel, Ajay, et al.
Published: (2024)
by: Patel, Ajay, et al.
Published: (2024)
Optimal Condition for Initialization Variance in Deep Neural Networks: An SGD Dynamics Perspective
by: Horii, Hiroshi, et al.
Published: (2025)
by: Horii, Hiroshi, et al.
Published: (2025)
Fishers for Free? Approximating the Fisher Information Matrix by Recycling the Squared Gradient Accumulator
by: Li, YuXin, et al.
Published: (2025)
by: Li, YuXin, et al.
Published: (2025)
IDInit: A Universal and Stable Initialization Method for Neural Network Training
by: Pan, Yu, et al.
Published: (2025)
by: Pan, Yu, et al.
Published: (2025)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
by: Yadav, Prateek, et al.
Published: (2023)
by: Yadav, Prateek, et al.
Published: (2023)
Rank Suggestion in Non-negative Matrix Factorization: Residual Sensitivity to Initial Conditions (RSIC)
by: Tunnell, Marc A., et al.
Published: (2024)
by: Tunnell, Marc A., et al.
Published: (2024)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
by: Pan, Bowen, et al.
Published: (2024)
by: Pan, Bowen, et al.
Published: (2024)
CISO: Species Distribution Modeling Conditioned on Incomplete Species Observations
by: Abdelwahed, Hager Radi, et al.
Published: (2025)
by: Abdelwahed, Hager Radi, et al.
Published: (2025)
Neural Network Training Techniques Regularize Optimization Trajectory: An Empirical Study
by: Chen, Cheng, et al.
Published: (2020)
by: Chen, Cheng, et al.
Published: (2020)
Localized, High-resolution Geographic Representations with Slepian Functions
by: Rao, Arjun, et al.
Published: (2026)
by: Rao, Arjun, et al.
Published: (2026)
MetaCluster: Enabling Deep Compression of Kolmogorov-Arnold Network
by: Raffel, Matthew, et al.
Published: (2025)
by: Raffel, Matthew, et al.
Published: (2025)
Interpretability of Graph Neural Networks to Assess Effects of Global Change Drivers on Ecological Networks
by: Anakok, Emre, et al.
Published: (2025)
by: Anakok, Emre, et al.
Published: (2025)
WTTFNet: A Weather-Time-Trajectory Fusion Network for Pedestrian Trajectory Prediction in Urban Complex
by: Wu, Ho Chun, et al.
Published: (2024)
by: Wu, Ho Chun, et al.
Published: (2024)
Principal Components for Neural Network Initialization
by: Phan, Nhan, et al.
Published: (2025)
by: Phan, Nhan, et al.
Published: (2025)
Deep Neural Network Initialization with Sparsity Inducing Activations
by: Price, Ilan, et al.
Published: (2024)
by: Price, Ilan, et al.
Published: (2024)
LION-DG: Layer-Informed Initialization with Deep Gradient Protocols for Accelerated Neural Network Training
by: Kim, Hyunjun
Published: (2026)
by: Kim, Hyunjun
Published: (2026)
Prior-Informed Neural Network Initialization: A Spectral Approach for Function Parameterizing Architectures
by: Torres, David Orlando Salazar, et al.
Published: (2026)
by: Torres, David Orlando Salazar, et al.
Published: (2026)
Text-to-Model: Text-Conditioned Neural Network Diffusion for Train-Once-for-All Personalization
by: Li, Zexi, et al.
Published: (2024)
by: Li, Zexi, et al.
Published: (2024)
FlashKAT: Understanding and Addressing Performance Bottlenecks in the Kolmogorov-Arnold Transformer
by: Raffel, Matthew, et al.
Published: (2025)
by: Raffel, Matthew, et al.
Published: (2025)
Fast Training of Sinusoidal Neural Fields via Scaling Initialization
by: Yeom, Taesun, et al.
Published: (2024)
by: Yeom, Taesun, et al.
Published: (2024)
Deep Fusion: Efficient Network Training via Pre-trained Initializations
by: Mazzawi, Hanna, et al.
Published: (2023)
by: Mazzawi, Hanna, et al.
Published: (2023)
CAWI: Copula-Aligned Weight Initialization for Randomized Neural Networks
by: Akhtar, Mushir, et al.
Published: (2026)
by: Akhtar, Mushir, et al.
Published: (2026)
Deconstructing the Goldilocks Zone of Neural Network Initialization
by: Vysogorets, Artem, et al.
Published: (2024)
by: Vysogorets, Artem, et al.
Published: (2024)
Model Merging via Data-Free Covariance Estimation
by: Hameed, Marawan Gamal Abdel, et al.
Published: (2026)
by: Hameed, Marawan Gamal Abdel, et al.
Published: (2026)
Deep Neural Network Training as Random Effects: An Optimization-Inference Duality
by: Yao, Minhao, et al.
Published: (2026)
by: Yao, Minhao, et al.
Published: (2026)
Similar Items
-
TokSuite: Measuring the Impact of Tokenizer Choice on Language Model Behavior
by: Altıntaş, Gül Sena, et al.
Published: (2025) -
Dataset Difficulty and the Role of Inductive Bias
by: Kwok, Devin, et al.
Published: (2024) -
Position: The Most Expensive Part of an LLM should be its Training Data
by: Kandpal, Nikhil, et al.
Published: (2025) -
Uncovering Model Processing Strategies with Non-Negative Per-Example Fisher Factorization
by: Matena, Michael, et al.
Published: (2023) -
Simultaneous linear connectivity of neural networks modulo permutation
by: Sharma, Ekansh, et al.
Published: (2024)