Dynamical Mean-Field Theory of Self-Attention Neural Networks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Poc-López, Ángel, Aguilera, Miguel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Precise Dynamics of Diagonal Linear Networks: A Unifying Analysis by Dynamical Mean-Field Theory
von: Nishiyama, Sota, et al.
Veröffentlicht: (2025)
von: Nishiyama, Sota, et al.
Veröffentlicht: (2025)
High-Dimensional Limit of Stochastic Gradient Flow via Dynamical Mean-Field Theory
von: Nishiyama, Sota, et al.
Veröffentlicht: (2026)
von: Nishiyama, Sota, et al.
Veröffentlicht: (2026)
Dynamic Mean-Field Theory for Continuous Random Networks
von: Zúñiga-Galindo, W. A.
Veröffentlicht: (2024)
von: Zúñiga-Galindo, W. A.
Veröffentlicht: (2024)
Growing Neural Networks: Dynamic Evolution through Gradient Descent
von: Radhakrishnan, Anil, et al.
Veröffentlicht: (2025)
von: Radhakrishnan, Anil, et al.
Veröffentlicht: (2025)
Recurrent Self-Attention Dynamics: An Energy-Agnostic Perspective from Jacobians
von: Tomihari, Akiyoshi, et al.
Veröffentlicht: (2025)
von: Tomihari, Akiyoshi, et al.
Veröffentlicht: (2025)
Controlled Langevin Dynamics for Sampling of Feedforward Neural Networks Trained with Minibatches
von: Zambon, Alessandro, et al.
Veröffentlicht: (2026)
von: Zambon, Alessandro, et al.
Veröffentlicht: (2026)
Formation of Representations in Neural Networks
von: Ziyin, Liu, et al.
Veröffentlicht: (2024)
von: Ziyin, Liu, et al.
Veröffentlicht: (2024)
A Theory of Saddle Escape in Deep Nonlinear Networks
von: Rawal, Divit, et al.
Veröffentlicht: (2026)
von: Rawal, Divit, et al.
Veröffentlicht: (2026)
A Dynamical Model of Neural Scaling Laws
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)
Asymptotics of SGD in Sequence-Single Index Models and Single-Layer Attention Networks
von: Arnaboldi, Luca, et al.
Veröffentlicht: (2025)
von: Arnaboldi, Luca, et al.
Veröffentlicht: (2025)
Graph Neural Networks Do Not Always Oversmooth
von: Epping, Bastian, et al.
Veröffentlicht: (2024)
von: Epping, Bastian, et al.
Veröffentlicht: (2024)
Topological Effects in Neural Network Field Theory
von: Ferko, Christian, et al.
Veröffentlicht: (2026)
von: Ferko, Christian, et al.
Veröffentlicht: (2026)
Deep Learning as Neural Low-Degree Filtering: A Spectral Theory of Hierarchical Feature Learning
von: Dandi, Yatin, et al.
Veröffentlicht: (2026)
von: Dandi, Yatin, et al.
Veröffentlicht: (2026)
Demolition and Reinforcement of Memories in Spin-Glass-like Neural Networks
von: Ventura, Enrico
Veröffentlicht: (2024)
von: Ventura, Enrico
Veröffentlicht: (2024)
The Rules-and-Facts Model for Simultaneous Generalization and Memorization in Neural Networks
von: Farné, Gabriele, et al.
Veröffentlicht: (2026)
von: Farné, Gabriele, et al.
Veröffentlicht: (2026)
Benchmarking Graph Neural Networks in Solving Hard Constraint Satisfaction Problems
von: Skenderi, Geri, et al.
Veröffentlicht: (2026)
von: Skenderi, Geri, et al.
Veröffentlicht: (2026)
A Random-Matrix Criterion for Initializing Gated Recurrent Neural Networks
von: Fioratti, Tommaso, et al.
Veröffentlicht: (2026)
von: Fioratti, Tommaso, et al.
Veröffentlicht: (2026)
A Federated Many-to-One Hopfield model for associative Neural Networks
von: Alessandrelli, Andrea, et al.
Veröffentlicht: (2026)
von: Alessandrelli, Andrea, et al.
Veröffentlicht: (2026)
Bayesian RG Flow in Neural Network Field Theories
von: Howard, Jessica N., et al.
Veröffentlicht: (2024)
von: Howard, Jessica N., et al.
Veröffentlicht: (2024)
Dynamical Decoupling of Generalization and Overfitting in Large Two-Layer Networks
von: Montanari, Andrea, et al.
Veröffentlicht: (2025)
von: Montanari, Andrea, et al.
Veröffentlicht: (2025)
Approximation Theory for Neural Networks: Old and New
von: Mukherjee, Soumendu Sundar, et al.
Veröffentlicht: (2026)
von: Mukherjee, Soumendu Sundar, et al.
Veröffentlicht: (2026)
Kernel Renormalization in Bayesian Deep Neural Networks: the Equivalent Wishart Ansatz in the Proportional Regime
von: Baglioni, Paolo, et al.
Veröffentlicht: (2026)
von: Baglioni, Paolo, et al.
Veröffentlicht: (2026)
Graph Neural Network Approach to Predicting Magnetization in Quasi-One-Dimensional Ising Systems
von: Slavin, V., et al.
Veröffentlicht: (2025)
von: Slavin, V., et al.
Veröffentlicht: (2025)
Siamese Neural Network for Label-Efficient Critical Phenomena Prediction in 3D Percolation Models
von: Wang, Shanshan, et al.
Veröffentlicht: (2025)
von: Wang, Shanshan, et al.
Veröffentlicht: (2025)
Statistical Physics of Deep Neural Networks: Generalization Capability, Beyond the Infinite Width, and Feature Learning
von: Ariosto, Sebastiano
Veröffentlicht: (2025)
von: Ariosto, Sebastiano
Veröffentlicht: (2025)
Neural Network Quantum Field Theory from Transformer Architectures
von: Ageev, Dmitry S., et al.
Veröffentlicht: (2026)
von: Ageev, Dmitry S., et al.
Veröffentlicht: (2026)
Dynamically Learning to Integrate in Recurrent Neural Networks
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
Dynamical Learning in Deep Asymmetric Recurrent Neural Networks
von: Badalotti, Davide, et al.
Veröffentlicht: (2025)
von: Badalotti, Davide, et al.
Veröffentlicht: (2025)
Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
Statistical Advantage of Softmax Attention: Insights from Single-Location Regression
von: Duranthon, O., et al.
Veröffentlicht: (2025)
von: Duranthon, O., et al.
Veröffentlicht: (2025)
The Quantization Model of Neural Scaling
von: Michaud, Eric J., et al.
Veröffentlicht: (2023)
von: Michaud, Eric J., et al.
Veröffentlicht: (2023)
Explaining Neural Scaling Laws
von: Bahri, Yasaman, et al.
Veröffentlicht: (2021)
von: Bahri, Yasaman, et al.
Veröffentlicht: (2021)
Gaussian Universality in Neural Network Dynamics with Generalized Structured Input Distributions
von: Bae, Jaeyong, et al.
Veröffentlicht: (2024)
von: Bae, Jaeyong, et al.
Veröffentlicht: (2024)
Field theory for optimal signal propagation in ResNets
von: Fischer, Kirsten, et al.
Veröffentlicht: (2023)
von: Fischer, Kirsten, et al.
Veröffentlicht: (2023)
Theory of Speciation Transitions in Diffusion Models with General Class Structure
von: Achilli, Beatrice, et al.
Veröffentlicht: (2026)
von: Achilli, Beatrice, et al.
Veröffentlicht: (2026)
Applications of Statistical Field Theory in Deep Learning
von: Ringel, Zohar, et al.
Veröffentlicht: (2025)
von: Ringel, Zohar, et al.
Veröffentlicht: (2025)
Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
Neural Scaling Laws Rooted in the Data Distribution
von: Brill, Ari
Veröffentlicht: (2024)
von: Brill, Ari
Veröffentlicht: (2024)
From Kernels to Features: A Multi-Scale Adaptive Theory of Feature Learning
von: Rubin, Noa, et al.
Veröffentlicht: (2025)
von: Rubin, Noa, et al.
Veröffentlicht: (2025)
Dissecting the Interplay of Attention Paths in a Statistical Mechanics Theory of Transformers
von: Tiberi, Lorenzo, et al.
Veröffentlicht: (2024)
von: Tiberi, Lorenzo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Precise Dynamics of Diagonal Linear Networks: A Unifying Analysis by Dynamical Mean-Field Theory
von: Nishiyama, Sota, et al.
Veröffentlicht: (2025) -
High-Dimensional Limit of Stochastic Gradient Flow via Dynamical Mean-Field Theory
von: Nishiyama, Sota, et al.
Veröffentlicht: (2026) -
Dynamic Mean-Field Theory for Continuous Random Networks
von: Zúñiga-Galindo, W. A.
Veröffentlicht: (2024) -
Growing Neural Networks: Dynamic Evolution through Gradient Descent
von: Radhakrishnan, Anil, et al.
Veröffentlicht: (2025) -
Recurrent Self-Attention Dynamics: An Energy-Agnostic Perspective from Jacobians
von: Tomihari, Akiyoshi, et al.
Veröffentlicht: (2025)