Geometric Dynamics of Signal Propagation Predict Trainability of Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Cowsik, Aditya, Nebabu, Tamra, Qi, Xiao-Liang, Ganguli, Surya |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Persian Rug: solving toy models of superposition using large-scale symmetries
by: Cowsik, Aditya, et al.
Published: (2024)
by: Cowsik, Aditya, et al.
Published: (2024)
Towards Distributed Neural Architectures
by: Cowsik, Aditya, et al.
Published: (2025)
by: Cowsik, Aditya, et al.
Published: (2025)
Solving adversarial examples requires solving exponential misalignment
by: Salvatore, Alessandro, et al.
Published: (2026)
by: Salvatore, Alessandro, et al.
Published: (2026)
An analytic theory of creativity in convolutional diffusion models
by: Kamb, Mason, et al.
Published: (2024)
by: Kamb, Mason, et al.
Published: (2024)
Infinite Limits of Multi-head Transformer Dynamics
by: Bordelon, Blake, et al.
Published: (2024)
by: Bordelon, Blake, et al.
Published: (2024)
Exact Learning Dynamics of In-Context Learning in Linear Transformers and Its Application to Non-Linear Transformers
by: Mainali, Nischal, et al.
Published: (2025)
by: Mainali, Nischal, et al.
Published: (2025)
Spread Complexity in Non-Hermitian Many-Body Localization Transition
by: Ganguli, Maitri
Published: (2024)
by: Ganguli, Maitri
Published: (2024)
Geometric Entropy and Retrieval Phase Transitions in Continuous Thermal Dense Associative Memory
by: Petrova, Tatiana, et al.
Published: (2026)
by: Petrova, Tatiana, et al.
Published: (2026)
Critical behavior of dirty parafermionic chains
by: Pandey, Akshat, et al.
Published: (2024)
by: Pandey, Akshat, et al.
Published: (2024)
Non-equilibrium active noise enhances generative memory in diffusion models
by: Behera, Agnish Kumar, et al.
Published: (2024)
by: Behera, Agnish Kumar, et al.
Published: (2024)
Reshaping Global Loop Structure to Accelerate Local Optimization by Smoothing Rugged Landscapes
by: Leleu, Timothee, et al.
Published: (2026)
by: Leleu, Timothee, et al.
Published: (2026)
Towards Understanding Inductive Bias in Transformers: A View From Infinity
by: Lavie, Itay, et al.
Published: (2024)
by: Lavie, Itay, et al.
Published: (2024)
Algorithmic Task Capture, Computational Complexity, and Inductive Bias of Infinite Transformers
by: Davidovich, Orit, et al.
Published: (2026)
by: Davidovich, Orit, et al.
Published: (2026)
Precise Dynamics of Diagonal Linear Networks: A Unifying Analysis by Dynamical Mean-Field Theory
by: Nishiyama, Sota, et al.
Published: (2025)
by: Nishiyama, Sota, et al.
Published: (2025)
Small Singular Values Matter: A Random Matrix Analysis of Transformer Models
by: Staats, Max, et al.
Published: (2024)
by: Staats, Max, et al.
Published: (2024)
Building Conformal Prediction Intervals with Approximate Message Passing
by: Clarté, Lucas, et al.
Published: (2024)
by: Clarté, Lucas, et al.
Published: (2024)
Dynamical Regimes of Multimodal Diffusion Models
by: Albrychiewicz, Emil, et al.
Published: (2026)
by: Albrychiewicz, Emil, et al.
Published: (2026)
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
by: Cagnetta, Francesco, et al.
Published: (2025)
by: Cagnetta, Francesco, et al.
Published: (2025)
A Dynamical Model of Neural Scaling Laws
by: Bordelon, Blake, et al.
Published: (2024)
by: Bordelon, Blake, et al.
Published: (2024)
Dynamics of Meta-learning Representation in the Teacher-student Scenario
by: Wang, Hui, et al.
Published: (2024)
by: Wang, Hui, et al.
Published: (2024)
Grokking as the Transition from Lazy to Rich Training Dynamics
by: Kumar, Tanishq, et al.
Published: (2023)
by: Kumar, Tanishq, et al.
Published: (2023)
Bias in Motion: Theoretical Insights into the Dynamics of Bias in SGD Training
by: Jain, Anchit, et al.
Published: (2024)
by: Jain, Anchit, et al.
Published: (2024)
Dynamical Mean-Field Theory of Self-Attention Neural Networks
by: Poc-López, Ángel, et al.
Published: (2024)
by: Poc-López, Ángel, et al.
Published: (2024)
The RL Perceptron: Generalisation Dynamics of Policy Learning in High Dimensions
by: Patel, Nishil, et al.
Published: (2023)
by: Patel, Nishil, et al.
Published: (2023)
The Interplay of Data Structure and Imbalance in the Learning Dynamics of Diffusion Models
by: Nicoletti, Flavio, et al.
Published: (2026)
by: Nicoletti, Flavio, et al.
Published: (2026)
Growing Neural Networks: Dynamic Evolution through Gradient Descent
by: Radhakrishnan, Anil, et al.
Published: (2025)
by: Radhakrishnan, Anil, et al.
Published: (2025)
Dynamical Decoupling of Generalization and Overfitting in Large Two-Layer Networks
by: Montanari, Andrea, et al.
Published: (2025)
by: Montanari, Andrea, et al.
Published: (2025)
Graph Neural Network Approach to Predicting Magnetization in Quasi-One-Dimensional Ising Systems
by: Slavin, V., et al.
Published: (2025)
by: Slavin, V., et al.
Published: (2025)
Training Dynamics of Nonlinear Contrastive Learning Model in the High Dimensional Limit
by: Meng, Lineghuan, et al.
Published: (2024)
by: Meng, Lineghuan, et al.
Published: (2024)
Disordered Dynamics in High Dimensions: Connections to Random Matrices and Machine Learning
by: Bordelon, Blake, et al.
Published: (2026)
by: Bordelon, Blake, et al.
Published: (2026)
Controlled Langevin Dynamics for Sampling of Feedforward Neural Networks Trained with Minibatches
by: Zambon, Alessandro, et al.
Published: (2026)
by: Zambon, Alessandro, et al.
Published: (2026)
Two-Point Deterministic Equivalence for Stochastic Gradient Dynamics in Linear Models
by: Atanasov, Alexander, et al.
Published: (2025)
by: Atanasov, Alexander, et al.
Published: (2025)
Siamese Neural Network for Label-Efficient Critical Phenomena Prediction in 3D Percolation Models
by: Wang, Shanshan, et al.
Published: (2025)
by: Wang, Shanshan, et al.
Published: (2025)
Why Diffusion Models Don't Memorize: The Role of Implicit Dynamical Regularization in Training
by: Bonnaire, Tony, et al.
Published: (2025)
by: Bonnaire, Tony, et al.
Published: (2025)
Stochastic Gradient Flow Dynamics of Test Risk and its Exact Solution for Weak Features
by: Veiga, Rodrigo, et al.
Published: (2024)
by: Veiga, Rodrigo, et al.
Published: (2024)
High-Dimensional Limit of Stochastic Gradient Flow via Dynamical Mean-Field Theory
by: Nishiyama, Sota, et al.
Published: (2026)
by: Nishiyama, Sota, et al.
Published: (2026)
Nature-Inspired Local Propagation
by: Betti, Alessandro, et al.
Published: (2024)
by: Betti, Alessandro, et al.
Published: (2024)
(How) Can Transformers Predict Pseudo-Random Numbers?
by: Tao, Tao, et al.
Published: (2025)
by: Tao, Tao, et al.
Published: (2025)
Transient learning dynamics drive escape from sharp valleys in Stochastic Gradient Descent
by: Yang, Ning, et al.
Published: (2026)
by: Yang, Ning, et al.
Published: (2026)
Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer
by: Bordelon, Blake, et al.
Published: (2025)
by: Bordelon, Blake, et al.
Published: (2025)
Similar Items
-
The Persian Rug: solving toy models of superposition using large-scale symmetries
by: Cowsik, Aditya, et al.
Published: (2024) -
Towards Distributed Neural Architectures
by: Cowsik, Aditya, et al.
Published: (2025) -
Solving adversarial examples requires solving exponential misalignment
by: Salvatore, Alessandro, et al.
Published: (2026) -
An analytic theory of creativity in convolutional diffusion models
by: Kamb, Mason, et al.
Published: (2024) -
Infinite Limits of Multi-head Transformer Dynamics
by: Bordelon, Blake, et al.
Published: (2024)