The emergence of clusters in self-attention dynamics
Fuente:
arXiv
Saved in:
| Main Authors: | Geshkovski, Borjan, Letrouit, Cyril, Polyanskiy, Yury, Rigollet, Philippe |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A mathematical perspective on Transformers
by: Geshkovski, Borjan, et al.
Published: (2023)
by: Geshkovski, Borjan, et al.
Published: (2023)
Dynamic metastability in the self-attention model
by: Geshkovski, Borjan, et al.
Published: (2024)
by: Geshkovski, Borjan, et al.
Published: (2024)
Synchronization of mean-field models on the circle
by: Polyanskiy, Yury, et al.
Published: (2025)
by: Polyanskiy, Yury, et al.
Published: (2025)
Quantitative Clustering in Mean-Field Transformer Models
by: Chen, Shi, et al.
Published: (2025)
by: Chen, Shi, et al.
Published: (2025)
Kinetic theory for Transformers and the lost-in-the-middle phenomenon
by: Duerinckx, Mitia, et al.
Published: (2026)
by: Duerinckx, Mitia, et al.
Published: (2026)
Constructive conditional normalizing flows
by: Geshkovski, Borjan, et al.
Published: (2026)
by: Geshkovski, Borjan, et al.
Published: (2026)
Clustering in Causal Attention Masking
by: Karagodin, Nikita, et al.
Published: (2024)
by: Karagodin, Nikita, et al.
Published: (2024)
Homogenized Transformers
by: Koubbi, Hugo, et al.
Published: (2026)
by: Koubbi, Hugo, et al.
Published: (2026)
On the number of modes of Gaussian kernel density estimators
by: Geshkovski, Borjan, et al.
Published: (2024)
by: Geshkovski, Borjan, et al.
Published: (2024)
Critical attention scaling in long-context transformers
by: Chen, Shi, et al.
Published: (2025)
by: Chen, Shi, et al.
Published: (2025)
Measure-to-measure interpolation using Transformers
by: Geshkovski, Borjan, et al.
Published: (2024)
by: Geshkovski, Borjan, et al.
Published: (2024)
Gluing methods for quantitative stability of optimal transport maps
by: Letrouit, Cyril, et al.
Published: (2024)
by: Letrouit, Cyril, et al.
Published: (2024)
Propagation of Chaos in Contextual Flow Maps
by: Chen, Shi, et al.
Published: (2026)
by: Chen, Shi, et al.
Published: (2026)
YuriiFormer: A Suite of Nesterov-Accelerated Transformers
by: Zimin, Aleksandr, et al.
Published: (2026)
by: Zimin, Aleksandr, et al.
Published: (2026)
Perceptrons and localization of attention's mean-field landscape
by: Álvarez-López, Antonio, et al.
Published: (2026)
by: Álvarez-López, Antonio, et al.
Published: (2026)
Measure-to-measure Regression with Transformers
by: Vandergrift, Matthew, et al.
Published: (2026)
by: Vandergrift, Matthew, et al.
Published: (2026)
Residual connections provably mitigate oversmoothing in graph neural networks
by: Chen, Ziang, et al.
Published: (2025)
by: Chen, Ziang, et al.
Published: (2025)
Birth-death dynamics for sampling: Global convergence, approximations and their asymptotics
by: Lu, Yulong, et al.
Published: (2022)
by: Lu, Yulong, et al.
Published: (2022)
Scaling Limits of Long-Context Transformers
by: Bruno, Giuseppe, et al.
Published: (2026)
by: Bruno, Giuseppe, et al.
Published: (2026)
Normalization in Attention Dynamics
by: Karagodin, Nikita, et al.
Published: (2025)
by: Karagodin, Nikita, et al.
Published: (2025)
Vector valued optimal transport: from dynamic to static formulations
by: Craig, Katy, et al.
Published: (2025)
by: Craig, Katy, et al.
Published: (2025)
On the Structure of Stationary Solutions to McKean-Vlasov Equations with Applications to Noisy Transformers
by: Balasubramanian, Krishnakumar, et al.
Published: (2025)
by: Balasubramanian, Krishnakumar, et al.
Published: (2025)
Learning functional components of PDEs from data using neural networks
by: Loman, Torkel E., et al.
Published: (2026)
by: Loman, Torkel E., et al.
Published: (2026)
Partial Differential Equations in the Age of Machine Learning: A Critical Synthesis of Classical, Machine Learning, and Hybrid Methods
by: Nooraiepour, Mohammad, et al.
Published: (2026)
by: Nooraiepour, Mohammad, et al.
Published: (2026)
Interaction-Force Transport Gradient Flows
by: Gladin, Egor, et al.
Published: (2024)
by: Gladin, Egor, et al.
Published: (2024)
Minimax Rates for the Estimation of Eigenpairs of Weighted Laplace-Beltrami Operators on Manifolds
by: Trillos, Nicolás García, et al.
Published: (2025)
by: Trillos, Nicolás García, et al.
Published: (2025)
Generalization Error Bounds for Picard-Type Operator Learning in Nonlinear Parabolic PDEs
by: Taniguchi, Koichi, et al.
Published: (2026)
by: Taniguchi, Koichi, et al.
Published: (2026)
Is Zero-Shot Super-Resolution Possible in Operator Learning?
by: Subedi, Unique, et al.
Published: (2026)
by: Subedi, Unique, et al.
Published: (2026)
Efficient Numerical Wave Propagation Enhanced By An End-to-End Deep Learning Model
by: Kaiser, Luis, et al.
Published: (2024)
by: Kaiser, Luis, et al.
Published: (2024)
Nonlocal Attention Operator: Materializing Hidden Knowledge Towards Interpretable Physics Discovery
by: Yu, Yue, et al.
Published: (2024)
by: Yu, Yue, et al.
Published: (2024)
On Probabilistic Embeddings in Optimal Dimension Reduction
by: Murray, Ryan, et al.
Published: (2024)
by: Murray, Ryan, et al.
Published: (2024)
Global Well-posedness and Convergence Analysis of Score-based Generative Models via Sharp Lipschitz Estimates
by: Mooney, Connor, et al.
Published: (2024)
by: Mooney, Connor, et al.
Published: (2024)
Learning PDE Solvers with Physics and Data: A Unifying View of Physics-Informed Neural Networks and Neural Operators
by: Dai, Yilong, et al.
Published: (2026)
by: Dai, Yilong, et al.
Published: (2026)
Adapting Noise to Data: Generative Flows from 1D Processes
by: Chemseddine, Jannis, et al.
Published: (2025)
by: Chemseddine, Jannis, et al.
Published: (2025)
Mean-Field Analysis for Learning Subspace-Sparse Polynomials with Gaussian Input
by: Chen, Ziang, et al.
Published: (2024)
by: Chen, Ziang, et al.
Published: (2024)
Improved Graph-based semi-supervised learning Schemes
by: Bozorgnia, Farid
Published: (2024)
by: Bozorgnia, Farid
Published: (2024)
Solving the Poisson Equation with Dirichlet data by shallow ReLU$^α$-networks: A regularity and approximation perspective
by: Vaishampayan, Malhar, et al.
Published: (2024)
by: Vaishampayan, Malhar, et al.
Published: (2024)
Double Coupling Architecture and Training Method for Optimization Problems of Differential Algebraic Equations with Parameters
by: Yang, Wenqiang, et al.
Published: (2026)
by: Yang, Wenqiang, et al.
Published: (2026)
Kernel Approximation of Fisher-Rao Gradient Flows
by: Zhu, Jia-Jie, et al.
Published: (2024)
by: Zhu, Jia-Jie, et al.
Published: (2024)
Deep Learning-Enhanced Calibration of the Heston Model: A Unified Framework
by: Zadgar, Arman, et al.
Published: (2025)
by: Zadgar, Arman, et al.
Published: (2025)
Similar Items
-
A mathematical perspective on Transformers
by: Geshkovski, Borjan, et al.
Published: (2023) -
Dynamic metastability in the self-attention model
by: Geshkovski, Borjan, et al.
Published: (2024) -
Synchronization of mean-field models on the circle
by: Polyanskiy, Yury, et al.
Published: (2025) -
Quantitative Clustering in Mean-Field Transformer Models
by: Chen, Shi, et al.
Published: (2025) -
Kinetic theory for Transformers and the lost-in-the-middle phenomenon
by: Duerinckx, Mitia, et al.
Published: (2026)