YuriiFormer: A Suite of Nesterov-Accelerated Transformers
Fuente:
arXiv
Guardado en:
| Autores principales: | Zimin, Aleksandr, Polyanskiy, Yury, Rigollet, Philippe |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Synchronization of mean-field models on the circle
por: Polyanskiy, Yury, et al.
Publicado: (2025)
por: Polyanskiy, Yury, et al.
Publicado: (2025)
Provable Acceleration of Nesterov's Accelerated Gradient Method over Heavy Ball Method in Training Over-Parameterized Neural Networks
por: Liu, Xin, et al.
Publicado: (2022)
por: Liu, Xin, et al.
Publicado: (2022)
Measure-to-measure interpolation using Transformers
por: Geshkovski, Borjan, et al.
Publicado: (2024)
por: Geshkovski, Borjan, et al.
Publicado: (2024)
Splat Regression Models
por: Daniels, Mara, et al.
Publicado: (2025)
por: Daniels, Mara, et al.
Publicado: (2025)
EMA-Nesterov: Stabilizing Nesterov's Lookahead for Accelerated Deep Learning Optimization
por: Yau, Chung-Yiu, et al.
Publicado: (2026)
por: Yau, Chung-Yiu, et al.
Publicado: (2026)
The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
por: Wang, Jinbo, et al.
Publicado: (2025)
por: Wang, Jinbo, et al.
Publicado: (2025)
Scaling Limits of Long-Context Transformers
por: Bruno, Giuseppe, et al.
Publicado: (2026)
por: Bruno, Giuseppe, et al.
Publicado: (2026)
Hierarchical Mixture-of-Experts with Two-Stage Optimization
por: Molodtsov, Gleb, et al.
Publicado: (2026)
por: Molodtsov, Gleb, et al.
Publicado: (2026)
A Concise Lyapunov Analysis of Nesterov's Accelerated Gradient Method
por: Liu, Jun
Publicado: (2025)
por: Liu, Jun
Publicado: (2025)
Functional multi-armed bandit and the best function identification problems
por: Dorn, Yuriy, et al.
Publicado: (2025)
por: Dorn, Yuriy, et al.
Publicado: (2025)
Nesterov Acceleration for Ensemble Kalman Inversion and Variants
por: Vernon, Sydney, et al.
Publicado: (2025)
por: Vernon, Sydney, et al.
Publicado: (2025)
Generalized Continuous-Time Models for Nesterov's Accelerated Gradient Methods
por: Park, Chanwoong, et al.
Publicado: (2024)
por: Park, Chanwoong, et al.
Publicado: (2024)
Provable Acceleration of Nesterov's Accelerated Gradient for Rectangular Matrix Factorization and Linear Neural Networks
por: Xu, Zhenghao, et al.
Publicado: (2024)
por: Xu, Zhenghao, et al.
Publicado: (2024)
Bench-MFG: A Benchmark Suite for Learning in Stationary Mean Field Games
por: Magnino, Lorenzo, et al.
Publicado: (2026)
por: Magnino, Lorenzo, et al.
Publicado: (2026)
Inference of Online Newton Methods with Nesterov's Accelerated Sketching
por: Wang, Haoxuan, et al.
Publicado: (2026)
por: Wang, Haoxuan, et al.
Publicado: (2026)
Accelerating RLHF Training with Reward Variance Increase
por: Yang, Zonglin, et al.
Publicado: (2025)
por: Yang, Zonglin, et al.
Publicado: (2025)
Provable Acceleration for Diffusion Models under Minimal Assumptions
por: Li, Gen, et al.
Publicado: (2024)
por: Li, Gen, et al.
Publicado: (2024)
Critical attention scaling in long-context transformers
por: Chen, Shi, et al.
Publicado: (2025)
por: Chen, Shi, et al.
Publicado: (2025)
A multilevel approach to accelerate the training of Transformers
por: Lauga, Guillaume, et al.
Publicado: (2025)
por: Lauga, Guillaume, et al.
Publicado: (2025)
PID Accelerated Temporal Difference Algorithms
por: Bedaywi, Mark, et al.
Publicado: (2024)
por: Bedaywi, Mark, et al.
Publicado: (2024)
Enhancing Stochastic Gradient Descent: A Unified Framework and Novel Acceleration Methods for Faster Convergence
por: Deng, Yichuan, et al.
Publicado: (2024)
por: Deng, Yichuan, et al.
Publicado: (2024)
Accelerating Cutting-Plane Algorithms via Reinforcement Learning Surrogates
por: Mana, Kyle, et al.
Publicado: (2023)
por: Mana, Kyle, et al.
Publicado: (2023)
Provable Accelerated Convergence of Nesterov's Momentum for Deep ReLU Neural Networks
por: Liao, Fangshuo, et al.
Publicado: (2023)
por: Liao, Fangshuo, et al.
Publicado: (2023)
Cactus: Accelerating Auto-Regressive Decoding with Constrained Acceptance Speculative Sampling
por: Hao, Yongchang, et al.
Publicado: (2026)
por: Hao, Yongchang, et al.
Publicado: (2026)
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
por: Meterez, Alexandru, et al.
Publicado: (2025)
por: Meterez, Alexandru, et al.
Publicado: (2025)
Training Infinitely Deep and Wide Transformers
por: Barboni, Raphaël, et al.
Publicado: (2026)
por: Barboni, Raphaël, et al.
Publicado: (2026)
Stability of Transformers under Layer Normalization
por: Kan, Kelvin, et al.
Publicado: (2025)
por: Kan, Kelvin, et al.
Publicado: (2025)
Spatial Transformers for Radio Map Estimation
por: Viet, Pham Q., et al.
Publicado: (2024)
por: Viet, Pham Q., et al.
Publicado: (2024)
Finite-Time Analysis of Gradient Descent for Shallow Transformers
por: Arda, Enes, et al.
Publicado: (2026)
por: Arda, Enes, et al.
Publicado: (2026)
Agentic Transformers Provably Learn to Search via Reinforcement Learning
por: Yang, Tong, et al.
Publicado: (2026)
por: Yang, Tong, et al.
Publicado: (2026)
Optimal Control for Transformer Architectures: Enhancing Generalization, Robustness and Efficiency
por: Kan, Kelvin, et al.
Publicado: (2025)
por: Kan, Kelvin, et al.
Publicado: (2025)
How Well Can Transformers Emulate In-context Newton's Method?
por: Giannou, Angeliki, et al.
Publicado: (2024)
por: Giannou, Angeliki, et al.
Publicado: (2024)
Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization
por: Huang, Yu, et al.
Publicado: (2025)
por: Huang, Yu, et al.
Publicado: (2025)
One-Layer Transformer Provably Learns One-Nearest Neighbor In Context
por: Li, Zihao, et al.
Publicado: (2024)
por: Li, Zihao, et al.
Publicado: (2024)
One Model, Two Roles: Emergent Specialization in a Shared Recurrent Transformer
por: Shen, Jucheng, et al.
Publicado: (2026)
por: Shen, Jucheng, et al.
Publicado: (2026)
From Optimization to Prediction: Transformer-Based Path-Flow Estimation to the Traffic Assignment Problem
por: Ameli, Mostafa, et al.
Publicado: (2025)
por: Ameli, Mostafa, et al.
Publicado: (2025)
On the Hopf-Cole Transform for Control-affine Schrödinger Bridge
por: Teter, Alexis, et al.
Publicado: (2025)
por: Teter, Alexis, et al.
Publicado: (2025)
Clustering in Causal Attention Masking
por: Karagodin, Nikita, et al.
Publicado: (2024)
por: Karagodin, Nikita, et al.
Publicado: (2024)
Nesterov acceleration in benignly non-convex landscapes
por: Gupta, Kanan, et al.
Publicado: (2024)
por: Gupta, Kanan, et al.
Publicado: (2024)
Toward TransfORmers: Revolutionizing the Solution of Mixed Integer Programs with Transformers
por: Cooper, Joshua F., et al.
Publicado: (2024)
por: Cooper, Joshua F., et al.
Publicado: (2024)
Ejemplares similares
-
Synchronization of mean-field models on the circle
por: Polyanskiy, Yury, et al.
Publicado: (2025) -
Provable Acceleration of Nesterov's Accelerated Gradient Method over Heavy Ball Method in Training Over-Parameterized Neural Networks
por: Liu, Xin, et al.
Publicado: (2022) -
Measure-to-measure interpolation using Transformers
por: Geshkovski, Borjan, et al.
Publicado: (2024) -
Splat Regression Models
por: Daniels, Mara, et al.
Publicado: (2025) -
EMA-Nesterov: Stabilizing Nesterov's Lookahead for Accelerated Deep Learning Optimization
por: Yau, Chung-Yiu, et al.
Publicado: (2026)