Scaling Limits of Long-Context Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Bruno, Giuseppe, Chen, Shi, Lin, Zhengjiang, Polyanskiy, Yury, Rigollet, Philippe |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Propagation of Chaos in Contextual Flow Maps
by: Chen, Shi, et al.
Published: (2026)
by: Chen, Shi, et al.
Published: (2026)
Residual connections provably mitigate oversmoothing in graph neural networks
by: Chen, Ziang, et al.
Published: (2025)
by: Chen, Ziang, et al.
Published: (2025)
Critical attention scaling in long-context transformers
by: Chen, Shi, et al.
Published: (2025)
by: Chen, Shi, et al.
Published: (2025)
Quantitative Clustering in Mean-Field Transformer Models
by: Chen, Shi, et al.
Published: (2025)
by: Chen, Shi, et al.
Published: (2025)
YuriiFormer: A Suite of Nesterov-Accelerated Transformers
by: Zimin, Aleksandr, et al.
Published: (2026)
by: Zimin, Aleksandr, et al.
Published: (2026)
L$^2$M: Mutual Information Scaling Law for Long-Context Language Modeling
by: Chen, Zhuo, et al.
Published: (2025)
by: Chen, Zhuo, et al.
Published: (2025)
Tail-Aware Information-Theoretic Generalization for RLHF and SGLD
by: Zhang, Huiming, et al.
Published: (2026)
by: Zhang, Huiming, et al.
Published: (2026)
A Score-Based Density Formula, with Applications in Diffusion Generative Models
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
Limit theorems of Chatterjee's rank correlation
by: Lin, Zhexiao, et al.
Published: (2022)
by: Lin, Zhexiao, et al.
Published: (2022)
A Gapped Scale-Sensitive Dimension and Lower Bounds for Offset Rademacher Complexity
by: Jia, Zeyu, et al.
Published: (2025)
by: Jia, Zeyu, et al.
Published: (2025)
Thresholds for Reconstruction of Random Hypergraphs From Graph Projections
by: Bresler, Guy, et al.
Published: (2025)
by: Bresler, Guy, et al.
Published: (2025)
High-Rate Quantized Matrix Multiplication II
by: Ordentlich, Or, et al.
Published: (2026)
by: Ordentlich, Or, et al.
Published: (2026)
Huber-energy measure quantization
by: Turinici, Gabriel
Published: (2022)
by: Turinici, Gabriel
Published: (2022)
Nonparametric MLE for Gaussian Location Mixtures: Certified Computation and Generic Behavior
by: Polyanskiy, Yury, et al.
Published: (2025)
by: Polyanskiy, Yury, et al.
Published: (2025)
Scaling Laws from Sequential Feature Recovery: A Solvable Hierarchical Model
by: Wortsman-Zurich, Arie, et al.
Published: (2026)
by: Wortsman-Zurich, Arie, et al.
Published: (2026)
Generalization and Scaling Laws for Mixture-of-Experts Transformers
by: Mayaki, Mansour Zoubeirou a
Published: (2026)
by: Mayaki, Mansour Zoubeirou a
Published: (2026)
Rate of convergence of the smoothed empirical Wasserstein distance
by: Block, Adam, et al.
Published: (2022)
by: Block, Adam, et al.
Published: (2022)
The Mean-Field Dynamics of Transformers
by: Rigollet, Philippe
Published: (2025)
by: Rigollet, Philippe
Published: (2025)
Statistical Consistency of Discrete-to-Continuous Limits of Determinantal Point Processes
by: Jaquard, Hugo, et al.
Published: (2026)
by: Jaquard, Hugo, et al.
Published: (2026)
Clustering in Causal Attention Masking
by: Karagodin, Nikita, et al.
Published: (2024)
by: Karagodin, Nikita, et al.
Published: (2024)
A Neural Scaling Law from Lottery Ticket Ensembling
by: Liu, Ziming, et al.
Published: (2023)
by: Liu, Ziming, et al.
Published: (2023)
Optimal Quantization for Matrix Multiplication
by: Ordentlich, Or, et al.
Published: (2024)
by: Ordentlich, Or, et al.
Published: (2024)
Central Limit Theorem for Bayesian Neural Network trained with Variational Inference
by: Descours, Arnaud, et al.
Published: (2024)
by: Descours, Arnaud, et al.
Published: (2024)
Limit Theorems for Stochastic Gradient Descent in High-Dimensional Single-Layer Networks
by: Rangriz, Parsa
Published: (2025)
by: Rangriz, Parsa
Published: (2025)
Transformers Handle Endogeneity in In-Context Linear Regression
by: Liang, Haodong, et al.
Published: (2024)
by: Liang, Haodong, et al.
Published: (2024)
Transformers as Decision Makers: Provable In-Context Reinforcement Learning via Supervised Pretraining
by: Lin, Licong, et al.
Published: (2023)
by: Lin, Licong, et al.
Published: (2023)
Gaussian mixture layers for neural networks
by: Chewi, Sinho, et al.
Published: (2025)
by: Chewi, Sinho, et al.
Published: (2025)
On the number of modes of Gaussian kernel density estimators
by: Geshkovski, Borjan, et al.
Published: (2024)
by: Geshkovski, Borjan, et al.
Published: (2024)
Partial and Exact Recovery of a Random Hypergraph from its Graph Projection
by: Bresler, Guy, et al.
Published: (2025)
by: Bresler, Guy, et al.
Published: (2025)
Homogenized Transformers
by: Koubbi, Hugo, et al.
Published: (2026)
by: Koubbi, Hugo, et al.
Published: (2026)
Bayesian Inference with Deep Weakly Nonlinear Networks
by: Hanin, Boris, et al.
Published: (2024)
by: Hanin, Boris, et al.
Published: (2024)
Statistical Limits in Random Tensors with Multiple Correlated Spikes
by: Qi, Yang, et al.
Published: (2025)
by: Qi, Yang, et al.
Published: (2025)
Convergence of Dirichlet Forms for MCMC Optimal Scaling with Dependent Target Distributions on Large Graphs
by: Ning, Ning
Published: (2022)
by: Ning, Ning
Published: (2022)
Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits
by: Chen, Fan, et al.
Published: (2025)
by: Chen, Fan, et al.
Published: (2025)
Matrix Denoising with Doubly Heteroscedastic Noise: Fundamental Limits and Optimal Spectral Methods
by: Zhang, Yihan, et al.
Published: (2024)
by: Zhang, Yihan, et al.
Published: (2024)
Superposition unifies power-law training dynamics
by: Chen, Zixin Jessie, et al.
Published: (2026)
by: Chen, Zixin Jessie, et al.
Published: (2026)
Statistical optimal transport
by: Chewi, Sinho, et al.
Published: (2024)
by: Chewi, Sinho, et al.
Published: (2024)
Multivariate Gaussian Approximation for Random Forest via Region-based Stabilization
by: Shi, Zhaoyang, et al.
Published: (2024)
by: Shi, Zhaoyang, et al.
Published: (2024)
Revenue Maximization Under Sequential Price Competition Via The Estimation Of s-Concave Demand Functions
by: Bracale, Daniele, et al.
Published: (2025)
by: Bracale, Daniele, et al.
Published: (2025)
How Particle-System Random Batch Methods Enhance Graph Transformer: Memory Efficiency and Parallel Computing Strategy
by: Liu, Hanwen, et al.
Published: (2025)
by: Liu, Hanwen, et al.
Published: (2025)
Similar Items
-
Propagation of Chaos in Contextual Flow Maps
by: Chen, Shi, et al.
Published: (2026) -
Residual connections provably mitigate oversmoothing in graph neural networks
by: Chen, Ziang, et al.
Published: (2025) -
Critical attention scaling in long-context transformers
by: Chen, Shi, et al.
Published: (2025) -
Quantitative Clustering in Mean-Field Transformer Models
by: Chen, Shi, et al.
Published: (2025) -
YuriiFormer: A Suite of Nesterov-Accelerated Transformers
by: Zimin, Aleksandr, et al.
Published: (2026)