Universal Approximation of Mean-Field Models via Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Biswal, Shiba, Elamvazhuthi, Karthik, Sonthalia, Rishi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Solution of the Critical Dynamics of the Mean-Field Kob-Andersen Model
by: Perrupato, Gianmarco, et al.
Published: (2025)
by: Perrupato, Gianmarco, et al.
Published: (2025)
Thermodynamics-inspired Explanations of Artificial Intelligence
by: Mehdi, Shams, et al.
Published: (2022)
by: Mehdi, Shams, et al.
Published: (2022)
Dissecting a Small Artificial Neural Network
by: Yang, Xiguang, et al.
Published: (2025)
by: Yang, Xiguang, et al.
Published: (2025)
The committee machine: Computational to statistical gaps in learning a two-layers neural network
by: Aubin, Benjamin, et al.
Published: (2018)
by: Aubin, Benjamin, et al.
Published: (2018)
Interpreting the Synchronization Gap: The Hidden Mechanism Inside Diffusion Transformers
by: Albrychiewicz, Emil, et al.
Published: (2026)
by: Albrychiewicz, Emil, et al.
Published: (2026)
Rejection-free Glauber Monte Carlo for the 2D Random Field Ising Model via Hierarchical Probabilistic Counters
by: Cattaneo, Luca, et al.
Published: (2026)
by: Cattaneo, Luca, et al.
Published: (2026)
Dissecting the Interplay of Attention Paths in a Statistical Mechanics Theory of Transformers
by: Tiberi, Lorenzo, et al.
Published: (2024)
by: Tiberi, Lorenzo, et al.
Published: (2024)
Gaussian Universality in Neural Network Dynamics with Generalized Structured Input Distributions
by: Bae, Jaeyong, et al.
Published: (2024)
by: Bae, Jaeyong, et al.
Published: (2024)
Universal Scaling Laws of Absorbing Phase Transitions in Artificial Deep Neural Networks
by: Tamai, Keiichi, et al.
Published: (2023)
by: Tamai, Keiichi, et al.
Published: (2023)
Optimal Protocols for Continual Learning via Statistical Physics and Control Theory
by: Mori, Francesco, et al.
Published: (2024)
by: Mori, Francesco, et al.
Published: (2024)
Biased Generalization in Diffusion Models
by: Garnier-Brun, Jerome, et al.
Published: (2026)
by: Garnier-Brun, Jerome, et al.
Published: (2026)
Dynamical Mean-Field Theory of Complex Systems on Sparse Directed Networks
by: Metz, Fernando L.
Published: (2024)
by: Metz, Fernando L.
Published: (2024)
Quantum Annealing in SK Model Employing Suzuki-Kubo-deGennes Quantum Ising Mean Field Dynamics
by: Das, Soumyaditya, et al.
Published: (2025)
by: Das, Soumyaditya, et al.
Published: (2025)
Rare Event Analysis of Large Language Models
by: Dorman, Jake McAllister, et al.
Published: (2026)
by: Dorman, Jake McAllister, et al.
Published: (2026)
Message Passing Variational Autoregressive Network for Solving Intractable Ising Models
by: Ma, Qunlong, et al.
Published: (2024)
by: Ma, Qunlong, et al.
Published: (2024)
Explaining the effects of non-convergent sampling in the training of Energy-Based Models
by: Agoritsas, Elisabeth, et al.
Published: (2023)
by: Agoritsas, Elisabeth, et al.
Published: (2023)
An Analytical Characterization of Sloppiness in Neural Networks: Insights from Linear Models
by: Mao, Jialin, et al.
Published: (2025)
by: Mao, Jialin, et al.
Published: (2025)
Exact solution of Dynamical Mean-Field Theory for a linear system with annealed disorder
by: Ferraro, Francesco, et al.
Published: (2024)
by: Ferraro, Francesco, et al.
Published: (2024)
Dataset-learning duality and emergent criticality
by: Kukleva, Ekaterina, et al.
Published: (2024)
by: Kukleva, Ekaterina, et al.
Published: (2024)
Thermodynamics of bidirectional associative memories
by: Barra, Adriano, et al.
Published: (2022)
by: Barra, Adriano, et al.
Published: (2022)
How do Probabilistic Graphical Models and Graph Neural Networks Look at Network Data?
by: Lapenna, Michela, et al.
Published: (2025)
by: Lapenna, Michela, et al.
Published: (2025)
Learning-at-Criticality in Large Language Models for Quantum Field Theory and Beyond
by: Cai, Xiansheng, et al.
Published: (2025)
by: Cai, Xiansheng, et al.
Published: (2025)
Fundamental operating regimes, hyper-parameter fine-tuning and glassiness: towards an interpretable replica-theory for trained restricted Boltzmann machines
by: Fachechi, Alberto, et al.
Published: (2024)
by: Fachechi, Alberto, et al.
Published: (2024)
Discrete generative diffusion models without stochastic differential equations: a tensor network approach
by: Causer, Luke, et al.
Published: (2024)
by: Causer, Luke, et al.
Published: (2024)
Transfer Learning in $\ell_1$ Regularized Regression: Hyperparameter Selection Strategy based on Sharp Asymptotic Analysis
by: Okajima, Koki, et al.
Published: (2024)
by: Okajima, Koki, et al.
Published: (2024)
Dynamic neuron approach to deep neural networks: Decoupling neurons for renormalization group analysis
by: Lee, Donghee, et al.
Published: (2024)
by: Lee, Donghee, et al.
Published: (2024)
Generalization vs. Specialization under Concept Shift
by: Nguyen, Alex, et al.
Published: (2024)
by: Nguyen, Alex, et al.
Published: (2024)
Coding schemes in neural networks learning classification tasks
by: van Meegen, Alexander, et al.
Published: (2024)
by: van Meegen, Alexander, et al.
Published: (2024)
Nonequilbrium physics of generative diffusion models
by: Yu, Zhendong, et al.
Published: (2024)
by: Yu, Zhendong, et al.
Published: (2024)
A replica analysis of under-bagging
by: Takahashi, Takashi
Published: (2024)
by: Takahashi, Takashi
Published: (2024)
Spring-block theory of feature learning in deep neural networks
by: Shi, Cheng, et al.
Published: (2024)
by: Shi, Cheng, et al.
Published: (2024)
Fast training and sampling of Restricted Boltzmann Machines
by: Béreux, Nicolas, et al.
Published: (2024)
by: Béreux, Nicolas, et al.
Published: (2024)
Machine learning the Ising transition: A comparison between discriminative and generative approaches
by: Zhang, Difei, et al.
Published: (2024)
by: Zhang, Difei, et al.
Published: (2024)
The Role of the Time-Dependent Hessian in High-Dimensional Optimization
by: Bonnaire, Tony, et al.
Published: (2024)
by: Bonnaire, Tony, et al.
Published: (2024)
Cascade of phase transitions in the training of Energy-based models
by: Bachtis, Dimitrios, et al.
Published: (2024)
by: Bachtis, Dimitrios, et al.
Published: (2024)
Analytic theory of dropout regularization
by: Mori, Francesco, et al.
Published: (2025)
by: Mori, Francesco, et al.
Published: (2025)
The Copycat Perceptron: Smashing Barriers Through Collective Learning
by: Catania, Giovanni, et al.
Published: (2023)
by: Catania, Giovanni, et al.
Published: (2023)
The autoregressive neural network architecture of the Boltzmann distribution of pairwise interacting spins systems
by: Biazzo, Indaco
Published: (2023)
by: Biazzo, Indaco
Published: (2023)
Distinct mechanisms underlying in-context learning in transformers
by: Gibson, Cole, et al.
Published: (2026)
by: Gibson, Cole, et al.
Published: (2026)
Two failure modes of deep transformers and how to avoid them: a unified theory of signal propagation at initialisation
by: Giorlandino, Alessio, et al.
Published: (2025)
by: Giorlandino, Alessio, et al.
Published: (2025)
Similar Items
-
Solution of the Critical Dynamics of the Mean-Field Kob-Andersen Model
by: Perrupato, Gianmarco, et al.
Published: (2025) -
Thermodynamics-inspired Explanations of Artificial Intelligence
by: Mehdi, Shams, et al.
Published: (2022) -
Dissecting a Small Artificial Neural Network
by: Yang, Xiguang, et al.
Published: (2025) -
The committee machine: Computational to statistical gaps in learning a two-layers neural network
by: Aubin, Benjamin, et al.
Published: (2018) -
Interpreting the Synchronization Gap: The Hidden Mechanism Inside Diffusion Transformers
by: Albrychiewicz, Emil, et al.
Published: (2026)