A multilevel approach to accelerate the training of Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Lauga, Guillaume, Chaumette, Maël, Desainte-Maréville, Edgar, Lasalle, Étienne, Lebeurrier, Arthur |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Path-conditioned training: a principled way to rescale ReLU neural networks
by: Lebeurrier, Arthur, et al.
Published: (2026)
by: Lebeurrier, Arthur, et al.
Published: (2026)
A block-coordinate descent framework for non-convex composite optimization. Application to sparse precision matrix estimation
by: Lauga, Guillaume
Published: (2026)
by: Lauga, Guillaume
Published: (2026)
Multiresolution Adaptive Block-Coordinate Forward-Backward for Image Reconstruction
by: Desainte-Maréville, Edgar, et al.
Published: (2026)
by: Desainte-Maréville, Edgar, et al.
Published: (2026)
Proximal basin hopping: global optimization with guarantees
by: Lauga, Guillaume, et al.
Published: (2026)
by: Lauga, Guillaume, et al.
Published: (2026)
Demystifying Manifold Constraints in LLM Pre-training
by: An, Kang, et al.
Published: (2026)
by: An, Kang, et al.
Published: (2026)
How Does Critical Batch Size Scale in Pre-training?
by: Zhang, Hanlin, et al.
Published: (2024)
by: Zhang, Hanlin, et al.
Published: (2024)
YuriiFormer: A Suite of Nesterov-Accelerated Transformers
by: Zimin, Aleksandr, et al.
Published: (2026)
by: Zimin, Aleksandr, et al.
Published: (2026)
Efficient and provably convergent end-to-end training of deep neural networks with linear constraints
by: Yang, Zonglin, et al.
Published: (2026)
by: Yang, Zonglin, et al.
Published: (2026)
Stability of Transformers under Layer Normalization
by: Kan, Kelvin, et al.
Published: (2025)
by: Kan, Kelvin, et al.
Published: (2025)
Training Infinitely Deep and Wide Transformers
by: Barboni, Raphaël, et al.
Published: (2026)
by: Barboni, Raphaël, et al.
Published: (2026)
Spatial Transformers for Radio Map Estimation
by: Viet, Pham Q., et al.
Published: (2024)
by: Viet, Pham Q., et al.
Published: (2024)
Finite-Time Analysis of Gradient Descent for Shallow Transformers
by: Arda, Enes, et al.
Published: (2026)
by: Arda, Enes, et al.
Published: (2026)
Optimal Control for Transformer Architectures: Enhancing Generalization, Robustness and Efficiency
by: Kan, Kelvin, et al.
Published: (2025)
by: Kan, Kelvin, et al.
Published: (2025)
Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization
by: Huang, Yu, et al.
Published: (2025)
by: Huang, Yu, et al.
Published: (2025)
How Well Can Transformers Emulate In-context Newton's Method?
by: Giannou, Angeliki, et al.
Published: (2024)
by: Giannou, Angeliki, et al.
Published: (2024)
Agentic Transformers Provably Learn to Search via Reinforcement Learning
by: Yang, Tong, et al.
Published: (2026)
by: Yang, Tong, et al.
Published: (2026)
On the Convergence of Overparameterized Problems: Inherent Properties of the Compositional Structure of Neural Networks
by: de Oliveira, Arthur Castello Branco, et al.
Published: (2025)
by: de Oliveira, Arthur Castello Branco, et al.
Published: (2025)
A multilevel framework for accelerating uSARA in radio-interferometric imaging
by: Lauga, Guillaume, et al.
Published: (2024)
by: Lauga, Guillaume, et al.
Published: (2024)
The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
by: Wang, Jinbo, et al.
Published: (2025)
by: Wang, Jinbo, et al.
Published: (2025)
One-Layer Transformer Provably Learns One-Nearest Neighbor In Context
by: Li, Zihao, et al.
Published: (2024)
by: Li, Zihao, et al.
Published: (2024)
One Model, Two Roles: Emergent Specialization in a Shared Recurrent Transformer
by: Shen, Jucheng, et al.
Published: (2026)
by: Shen, Jucheng, et al.
Published: (2026)
Anderson acceleration for iteratively reweighted $\ell_1$ algorithm
by: Li, Kexin
Published: (2024)
by: Li, Kexin
Published: (2024)
From Optimization to Prediction: Transformer-Based Path-Flow Estimation to the Traffic Assignment Problem
by: Ameli, Mostafa, et al.
Published: (2025)
by: Ameli, Mostafa, et al.
Published: (2025)
On the Hopf-Cole Transform for Control-affine Schrödinger Bridge
by: Teter, Alexis, et al.
Published: (2025)
by: Teter, Alexis, et al.
Published: (2025)
gridfm-datakit-v1: A Python Library for Scalable and Realistic Power Flow and Optimal Power Flow Data Generation
by: Puech, Alban, et al.
Published: (2025)
by: Puech, Alban, et al.
Published: (2025)
Toward TransfORmers: Revolutionizing the Solution of Mixed Integer Programs with Transformers
by: Cooper, Joshua F., et al.
Published: (2024)
by: Cooper, Joshua F., et al.
Published: (2024)
Precise gradient descent training dynamics for finite-width multi-layer neural networks
by: Han, Qiyang, et al.
Published: (2025)
by: Han, Qiyang, et al.
Published: (2025)
Transformers as Support Vector Machines
by: Tarzanagh, Davoud Ataee, et al.
Published: (2023)
by: Tarzanagh, Davoud Ataee, et al.
Published: (2023)
A Minimalist Bayesian Framework for Stochastic Optimization
by: Wang, Kaizheng
Published: (2025)
by: Wang, Kaizheng
Published: (2025)
The Asymptotic Behavior of Attention in Transformers
by: Abella, Álvaro Rodríguez, et al.
Published: (2024)
by: Abella, Álvaro Rodríguez, et al.
Published: (2024)
A Novel Unified Parametric Assumption for Nonconvex Optimization
by: Riabinin, Artem, et al.
Published: (2025)
by: Riabinin, Artem, et al.
Published: (2025)
A Rod Flow Model for Adam at the Edge of Stability
by: Regis, Eric, et al.
Published: (2026)
by: Regis, Eric, et al.
Published: (2026)
Effective Frontiers: A Unification of Neural Scaling Laws
by: Zou, Jiaxuan, et al.
Published: (2026)
by: Zou, Jiaxuan, et al.
Published: (2026)
A Median Perspective on Unlabeled Data for Out-of-Distribution Detection
by: Abbas, Momin, et al.
Published: (2025)
by: Abbas, Momin, et al.
Published: (2025)
A Unified Framework for Gradient Aggregation in Multi-Objective Optimization
by: Hu, Zeou, et al.
Published: (2026)
by: Hu, Zeou, et al.
Published: (2026)
ARO: A New Lens On Matrix Optimization For Large Models
by: Gong, Wenbo, et al.
Published: (2026)
by: Gong, Wenbo, et al.
Published: (2026)
Stochastic Optimization with Constraints: A Non-asymptotic Instance-Dependent Analysis
by: Khamaru, Koulik
Published: (2024)
by: Khamaru, Koulik
Published: (2024)
A Convexity-dependent Two-Phase Training Algorithm for Deep Neural Networks
by: Hrycej, Tomas, et al.
Published: (2025)
by: Hrycej, Tomas, et al.
Published: (2025)
A second-order-like optimizer with adaptive gradient scaling for deep learning
by: Bolte, Jérôme, et al.
Published: (2024)
by: Bolte, Jérôme, et al.
Published: (2024)
A multiobjective continuation method to compute the regularization path of deep neural networks
by: Amakor, Augustina C., et al.
Published: (2023)
by: Amakor, Augustina C., et al.
Published: (2023)
Similar Items
-
Path-conditioned training: a principled way to rescale ReLU neural networks
by: Lebeurrier, Arthur, et al.
Published: (2026) -
A block-coordinate descent framework for non-convex composite optimization. Application to sparse precision matrix estimation
by: Lauga, Guillaume
Published: (2026) -
Multiresolution Adaptive Block-Coordinate Forward-Backward for Image Reconstruction
by: Desainte-Maréville, Edgar, et al.
Published: (2026) -
Proximal basin hopping: global optimization with guarantees
by: Lauga, Guillaume, et al.
Published: (2026) -
Demystifying Manifold Constraints in LLM Pre-training
by: An, Kang, et al.
Published: (2026)