A multiscale analysis of mean-field transformers in the moderate interaction regime
Fuente:
arXiv
Saved in:
| Main Authors: | Bruno, Giuseppe, Pasqualotto, Federico, Agazzi, Andrea |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Emergence of meta-stable clustering in mean-field transformer models
by: Bruno, Giuseppe, et al.
Published: (2024)
by: Bruno, Giuseppe, et al.
Published: (2024)
Stochastic Scaling Limits and Synchronization by Noise in Deep Transformer Models
by: Agazzi, Andrea, et al.
Published: (2026)
by: Agazzi, Andrea, et al.
Published: (2026)
Quantitative convergence of trained single layer neural networks to Gaussian processes
by: Mosig, Eloy, et al.
Published: (2025)
by: Mosig, Eloy, et al.
Published: (2025)
Temporal-difference learning with nonlinear function approximation: lazy training and mean field regimes
by: Agazzi, Andrea, et al.
Published: (2019)
by: Agazzi, Andrea, et al.
Published: (2019)
Function approximation by neural nets in the mean-field regime: Entropic regularization and controlled McKean-Vlasov dynamics
by: Tzen, Belinda, et al.
Published: (2020)
by: Tzen, Belinda, et al.
Published: (2020)
Mirror Descent-Ascent for mean-field min-max problems
by: Lascu, Razvan-Andrei, et al.
Published: (2024)
by: Lascu, Razvan-Andrei, et al.
Published: (2024)
A Fisher-Rao gradient flow for entropic mean-field min-max games
by: Lascu, Razvan-Andrei, et al.
Published: (2024)
by: Lascu, Razvan-Andrei, et al.
Published: (2024)
Non-convex entropic mean-field optimization via Best Response flow
by: Lascu, Razvan-Andrei, et al.
Published: (2025)
by: Lascu, Razvan-Andrei, et al.
Published: (2025)
On propagation of chaos for the Fisher-Rao gradient flow in entropic mean-field optimization
by: Lazić, Petra, et al.
Published: (2026)
by: Lazić, Petra, et al.
Published: (2026)
Low-degree Lower bounds for clustering in moderate dimension
by: Carpentier, Alexandra, et al.
Published: (2026)
by: Carpentier, Alexandra, et al.
Published: (2026)
Independent projections of diffusions: Gradient flows for variational inference and optimal mean field approximations
by: Lacker, Daniel
Published: (2023)
by: Lacker, Daniel
Published: (2023)
Non--exchangeable mean field games with moderate interactions and common noise
by: Djete, Mao Fabrice
Published: (2026)
by: Djete, Mao Fabrice
Published: (2026)
Noise-induced stabilization in a chemical reaction network without boundary effects
by: Agazzi, Andrea, et al.
Published: (2025)
by: Agazzi, Andrea, et al.
Published: (2025)
Large deviation principles for convolutional Bayesian neural networks
by: Bassetti, Federico, et al.
Published: (2026)
by: Bassetti, Federico, et al.
Published: (2026)
Data-driven approximation of transfer operators for mean-field stochastic differential equations
by: Ioannou, Eirini, et al.
Published: (2025)
by: Ioannou, Eirini, et al.
Published: (2025)
Fractional neural attention for efficient multiscale sequence processing
by: Qu, Cheng Kevin, et al.
Published: (2025)
by: Qu, Cheng Kevin, et al.
Published: (2025)
A simple algorithm for output range analysis for deep neural networks
by: Rojas, Helder, et al.
Published: (2024)
by: Rojas, Helder, et al.
Published: (2024)
Parameter uncertainties for imperfect surrogate models in the low-noise regime
by: Swinburne, Thomas D, et al.
Published: (2024)
by: Swinburne, Thomas D, et al.
Published: (2024)
On the distance between mean and geometric median in high dimensions
by: Schwank, Richard, et al.
Published: (2025)
by: Schwank, Richard, et al.
Published: (2025)
Asymptotically optimal sequential change detection for bounded means
by: Ram, Ashwin, et al.
Published: (2026)
by: Ram, Ashwin, et al.
Published: (2026)
Singular-limit analysis of gradient descent with noise injection
by: Shalova, Anna, et al.
Published: (2024)
by: Shalova, Anna, et al.
Published: (2024)
Scaling Limits of Long-Context Transformers
by: Bruno, Giuseppe, et al.
Published: (2026)
by: Bruno, Giuseppe, et al.
Published: (2026)
Effective continuous equations for adaptive SGD: a stochastic analysis view
by: Callisti, Luca, et al.
Published: (2025)
by: Callisti, Luca, et al.
Published: (2025)
Machine learning-based system reliability analysis with Gaussian Process Regression
by: Zhou, Lisang, et al.
Published: (2024)
by: Zhou, Lisang, et al.
Published: (2024)
Wasserstein Convergence of Score-based Generative Models under Semiconvexity and Discontinuous Gradients
by: Bruno, Stefano, et al.
Published: (2025)
by: Bruno, Stefano, et al.
Published: (2025)
Learning operators on labelled conditional distributions with applications to mean field control of non exchangeable systems
by: Mekkaoui, Samy, et al.
Published: (2026)
by: Mekkaoui, Samy, et al.
Published: (2026)
Weighted Random Dot Product Graphs
by: Marenco, Bernardo, et al.
Published: (2025)
by: Marenco, Bernardo, et al.
Published: (2025)
What can we learn from signals and systems in a transformer? Insights for probabilistic modeling and inference architecture
by: Chang, Heng-Sheng, et al.
Published: (2025)
by: Chang, Heng-Sheng, et al.
Published: (2025)
Change of measure through the Legendre transform
by: Picard-Weibel, Antoine, et al.
Published: (2022)
by: Picard-Weibel, Antoine, et al.
Published: (2022)
Scaling Laws from Sequential Feature Recovery: A Solvable Hierarchical Model
by: Wortsman-Zurich, Arie, et al.
Published: (2026)
by: Wortsman-Zurich, Arie, et al.
Published: (2026)
Propagation of chaos for first-order mean-field systems with non-attractive moderately singular interaction
by: Höfer, Richard M., et al.
Published: (2025)
by: Höfer, Richard M., et al.
Published: (2025)
Diffusion models for Gaussian distributions: Exact solutions and Wasserstein errors
by: Pierret, Emile, et al.
Published: (2024)
by: Pierret, Emile, et al.
Published: (2024)
Flatness-Aware Stochastic Gradient Langevin Dynamics
by: Bruno, Stefano, et al.
Published: (2025)
by: Bruno, Stefano, et al.
Published: (2025)
Gaussian random field approximation via Stein's method with applications to wide random neural networks
by: Balasubramanian, Krishnakumar, et al.
Published: (2023)
by: Balasubramanian, Krishnakumar, et al.
Published: (2023)
A Markovian Model for Learning-to-Optimize
by: Sucker, Michael, et al.
Published: (2024)
by: Sucker, Michael, et al.
Published: (2024)
A distance function for stochastic matrices
by: Lee, Antony R., et al.
Published: (2024)
by: Lee, Antony R., et al.
Published: (2024)
A characterization of sample adaptivity in UCB data
by: Chen, Yilun, et al.
Published: (2025)
by: Chen, Yilun, et al.
Published: (2025)
Accelerating Constrained Sampling: A Large Deviations Approach
by: Wang, Yingli, et al.
Published: (2025)
by: Wang, Yingli, et al.
Published: (2025)
A User's Guide to Sampling Strategies for Sliced Optimal Transport
by: Sisouk, Keanu, et al.
Published: (2025)
by: Sisouk, Keanu, et al.
Published: (2025)
Symmetries in Overparametrized Neural Networks: A Mean-Field View
by: Maass, Javier, et al.
Published: (2024)
by: Maass, Javier, et al.
Published: (2024)
Similar Items
-
Emergence of meta-stable clustering in mean-field transformer models
by: Bruno, Giuseppe, et al.
Published: (2024) -
Stochastic Scaling Limits and Synchronization by Noise in Deep Transformer Models
by: Agazzi, Andrea, et al.
Published: (2026) -
Quantitative convergence of trained single layer neural networks to Gaussian processes
by: Mosig, Eloy, et al.
Published: (2025) -
Temporal-difference learning with nonlinear function approximation: lazy training and mean field regimes
by: Agazzi, Andrea, et al.
Published: (2019) -
Function approximation by neural nets in the mean-field regime: Entropic regularization and controlled McKean-Vlasov dynamics
by: Tzen, Belinda, et al.
Published: (2020)