Stability of Transformers under Layer Normalization
Fuente:
arXiv
Guardado en:
| Autores principales: | Kan, Kelvin, Li, Xingjian, Zhang, Benjamin J., Sahai, Tuhin, Osher, Stanley, Kumar, Krishna, Katsoulakis, Markos A. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Optimal Control for Transformer Architectures: Enhancing Generalization, Robustness and Efficiency
por: Kan, Kelvin, et al.
Publicado: (2025)
por: Kan, Kelvin, et al.
Publicado: (2025)
Zero-Shot Transferable Solution Method for Parametric Optimal Control Problems
por: Li, Xingjian, et al.
Publicado: (2025)
por: Li, Xingjian, et al.
Publicado: (2025)
Dimension-Free Convergence of Discrete Diffusion Models: Adjoint Equations Induce the Right Space
por: Kan, Kelvin, et al.
Publicado: (2026)
por: Kan, Kelvin, et al.
Publicado: (2026)
OT-Transformer: A Continuous-time Transformer Architecture with Optimal Transport Regularization
por: Kan, Kelvin, et al.
Publicado: (2025)
por: Kan, Kelvin, et al.
Publicado: (2025)
Sparse Transformer Architectures via Regularized Wasserstein Proximal Operator with $L_1$ Prior
por: Han, Fuqun, et al.
Publicado: (2025)
por: Han, Fuqun, et al.
Publicado: (2025)
Thinking Out of the Box: Hybrid SAT Solving by Unconstrained Continuous Optimization
por: Zhang, Zhiwei, et al.
Publicado: (2025)
por: Zhang, Zhiwei, et al.
Publicado: (2025)
One-Layer Transformer Provably Learns One-Nearest Neighbor In Context
por: Li, Zihao, et al.
Publicado: (2024)
por: Li, Zihao, et al.
Publicado: (2024)
Proximal optimal transport divergences
por: Baptista, Ricardo, et al.
Publicado: (2025)
por: Baptista, Ricardo, et al.
Publicado: (2025)
A Unified Framework for Gradient Aggregation in Multi-Objective Optimization
por: Hu, Zeou, et al.
Publicado: (2026)
por: Hu, Zeou, et al.
Publicado: (2026)
Optimization and Generalization Guarantees for Weight Normalization
por: Cisneros-Velarde, Pedro, et al.
Publicado: (2024)
por: Cisneros-Velarde, Pedro, et al.
Publicado: (2024)
Global Convergence and Rich Feature Learning in $L$-Layer Infinite-Width Neural Networks under $μ$P Parametrization
por: Chen, Zixiang, et al.
Publicado: (2025)
por: Chen, Zixiang, et al.
Publicado: (2025)
A Natural Primal-Dual Hybrid Gradient Method for Adversarial Neural Network Training on Solving Partial Differential Equations
por: Liu, Shu, et al.
Publicado: (2024)
por: Liu, Shu, et al.
Publicado: (2024)
Provable Acceleration for Diffusion Models under Minimal Assumptions
por: Li, Gen, et al.
Publicado: (2024)
por: Li, Gen, et al.
Publicado: (2024)
Tree-Preconditioned Differentiable Optimization and Axioms as Layers
por: Liao, Yuexin
Publicado: (2025)
por: Liao, Yuexin
Publicado: (2025)
Parameter-Efficient Distributional RL via Normalizing Flows and a Geometry-Aware Cramér Surrogate
por: C., Simo Alami, et al.
Publicado: (2025)
por: C., Simo Alami, et al.
Publicado: (2025)
Differentiable Distributionally Robust Optimization Layers
por: Ma, Xutao, et al.
Publicado: (2024)
por: Ma, Xutao, et al.
Publicado: (2024)
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
por: Sheen, Heejune, et al.
Publicado: (2024)
por: Sheen, Heejune, et al.
Publicado: (2024)
SGD at the Edge of Stability: The Stochastic Sharpness Gap
por: Liao, Fangshuo, et al.
Publicado: (2026)
por: Liao, Fangshuo, et al.
Publicado: (2026)
A Rod Flow Model for Adam at the Edge of Stability
por: Regis, Eric, et al.
Publicado: (2026)
por: Regis, Eric, et al.
Publicado: (2026)
Uncovering Symmetry Transfer in Large Language Models via Layer-Peeled Optimization
por: Du, Zhehang, et al.
Publicado: (2026)
por: Du, Zhehang, et al.
Publicado: (2026)
Training Infinitely Deep and Wide Transformers
por: Barboni, Raphaël, et al.
Publicado: (2026)
por: Barboni, Raphaël, et al.
Publicado: (2026)
Spatial Transformers for Radio Map Estimation
por: Viet, Pham Q., et al.
Publicado: (2024)
por: Viet, Pham Q., et al.
Publicado: (2024)
Stochastic Subgradient Methods with Guaranteed Global Stability in Nonsmooth Nonconvex Optimization
por: Xiao, Nachuan, et al.
Publicado: (2023)
por: Xiao, Nachuan, et al.
Publicado: (2023)
DR-SAC: Distributionally Robust Soft Actor-Critic for Reinforcement Learning under Uncertainty
por: Cui, Mingxuan, et al.
Publicado: (2025)
por: Cui, Mingxuan, et al.
Publicado: (2025)
A multilevel approach to accelerate the training of Transformers
por: Lauga, Guillaume, et al.
Publicado: (2025)
por: Lauga, Guillaume, et al.
Publicado: (2025)
Rod Flow: A Continuous-Time Model for Gradient Descent at the Edge of Stability
por: Regis, Eric, et al.
Publicado: (2026)
por: Regis, Eric, et al.
Publicado: (2026)
YuriiFormer: A Suite of Nesterov-Accelerated Transformers
por: Zimin, Aleksandr, et al.
Publicado: (2026)
por: Zimin, Aleksandr, et al.
Publicado: (2026)
Finite-Time Analysis of Gradient Descent for Shallow Transformers
por: Arda, Enes, et al.
Publicado: (2026)
por: Arda, Enes, et al.
Publicado: (2026)
A Queueing-Theoretic Framework for Stability Analysis of LLM Inference with KV Cache Memory Constraints
por: Nie, Chengyi, et al.
Publicado: (2026)
por: Nie, Chengyi, et al.
Publicado: (2026)
How Well Can Transformers Emulate In-context Newton's Method?
por: Giannou, Angeliki, et al.
Publicado: (2024)
por: Giannou, Angeliki, et al.
Publicado: (2024)
Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization
por: Huang, Yu, et al.
Publicado: (2025)
por: Huang, Yu, et al.
Publicado: (2025)
Agentic Transformers Provably Learn to Search via Reinforcement Learning
por: Yang, Tong, et al.
Publicado: (2026)
por: Yang, Tong, et al.
Publicado: (2026)
Adaptively Robust LLM Inference Optimization under Prediction Uncertainty
por: Chen, Zixi, et al.
Publicado: (2025)
por: Chen, Zixi, et al.
Publicado: (2025)
The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
por: Wang, Jinbo, et al.
Publicado: (2025)
por: Wang, Jinbo, et al.
Publicado: (2025)
Reinforcement Learning under Latent Dynamics: Toward Statistical and Algorithmic Modularity
por: Amortila, Philip, et al.
Publicado: (2024)
por: Amortila, Philip, et al.
Publicado: (2024)
One Model, Two Roles: Emergent Specialization in a Shared Recurrent Transformer
por: Shen, Jucheng, et al.
Publicado: (2026)
por: Shen, Jucheng, et al.
Publicado: (2026)
From Optimization to Prediction: Transformer-Based Path-Flow Estimation to the Traffic Assignment Problem
por: Ameli, Mostafa, et al.
Publicado: (2025)
por: Ameli, Mostafa, et al.
Publicado: (2025)
On the Hopf-Cole Transform for Control-affine Schrödinger Bridge
por: Teter, Alexis, et al.
Publicado: (2025)
por: Teter, Alexis, et al.
Publicado: (2025)
Stability of Primal-Dual Gradient Flow Dynamics for Multi-Block Convex Optimization Problems
por: Ozaslan, Ibrahim K., et al.
Publicado: (2024)
por: Ozaslan, Ibrahim K., et al.
Publicado: (2024)
A Methodology Establishing Linear Convergence of Adaptive Gradient Methods under PL Inequality
por: Chakrabarti, Kushal, et al.
Publicado: (2024)
por: Chakrabarti, Kushal, et al.
Publicado: (2024)
Ejemplares similares
-
Optimal Control for Transformer Architectures: Enhancing Generalization, Robustness and Efficiency
por: Kan, Kelvin, et al.
Publicado: (2025) -
Zero-Shot Transferable Solution Method for Parametric Optimal Control Problems
por: Li, Xingjian, et al.
Publicado: (2025) -
Dimension-Free Convergence of Discrete Diffusion Models: Adjoint Equations Induce the Right Space
por: Kan, Kelvin, et al.
Publicado: (2026) -
OT-Transformer: A Continuous-time Transformer Architecture with Optimal Transport Regularization
por: Kan, Kelvin, et al.
Publicado: (2025) -
Sparse Transformer Architectures via Regularized Wasserstein Proximal Operator with $L_1$ Prior
por: Han, Fuqun, et al.
Publicado: (2025)