Linear attention is (maybe) all you need (to understand transformer optimization)
Fuente:
arXiv
Salvato in:
| Autori principali: | Ahn, Kwangjun, Cheng, Xiang, Song, Minhak, Yun, Chulhee, Jadbabaie, Ali, Sra, Suvrit |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
How to escape sharp minima with random perturbations
di: Ahn, Kwangjun, et al.
Pubblicazione: (2023)
di: Ahn, Kwangjun, et al.
Pubblicazione: (2023)
Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training
di: Song, Minhak, et al.
Pubblicazione: (2025)
di: Song, Minhak, et al.
Pubblicazione: (2025)
Does SGD really happen in tiny subspaces?
di: Song, Minhak, et al.
Pubblicazione: (2024)
di: Song, Minhak, et al.
Pubblicazione: (2024)
Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch Regime
di: Baek, Beomhan, et al.
Pubblicazione: (2025)
di: Baek, Beomhan, et al.
Pubblicazione: (2025)
Riemannian Bilevel Optimization
di: Dutta, Sanchayan, et al.
Pubblicazione: (2024)
di: Dutta, Sanchayan, et al.
Pubblicazione: (2024)
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
di: Kim, Jihwan, et al.
Pubblicazione: (2026)
di: Kim, Jihwan, et al.
Pubblicazione: (2026)
Toward generalizable learning of all (linear) first-order methods via memory augmented Transformers
di: Dutta, Sanchayan, et al.
Pubblicazione: (2024)
di: Dutta, Sanchayan, et al.
Pubblicazione: (2024)
Linearly Convergent Algorithms for Nonsmooth Problems with Unknown Smooth Pieces
di: Zhang, Zhe, et al.
Pubblicazione: (2025)
di: Zhang, Zhe, et al.
Pubblicazione: (2025)
Cost-Driven Representation Learning for Linear Quadratic Gaussian Control: Part I
di: Tian, Yi, et al.
Pubblicazione: (2022)
di: Tian, Yi, et al.
Pubblicazione: (2022)
Cost-Driven Representation Learning for Linear Quadratic Gaussian Control: Part II
di: Tian, Yi, et al.
Pubblicazione: (2026)
di: Tian, Yi, et al.
Pubblicazione: (2026)
Adam with model exponential moving average is effective for nonconvex optimization
di: Ahn, Kwangjun, et al.
Pubblicazione: (2024)
di: Ahn, Kwangjun, et al.
Pubblicazione: (2024)
Implicit Bias in Matrix Factorization and its Explicit Realization in a New Architecture
di: Hou, Yikun, et al.
Pubblicazione: (2025)
di: Hou, Yikun, et al.
Pubblicazione: (2025)
First-Order Methods for Linearly Constrained Bilevel Optimization
di: Kornowski, Guy, et al.
Pubblicazione: (2024)
di: Kornowski, Guy, et al.
Pubblicazione: (2024)
Dion: Distributed Orthonormalized Updates
di: Ahn, Kwangjun, et al.
Pubblicazione: (2025)
di: Ahn, Kwangjun, et al.
Pubblicazione: (2025)
General framework for online-to-nonconvex conversion: Schedule-free SGD is also effective for nonconvex optimization
di: Ahn, Kwangjun, et al.
Pubblicazione: (2024)
di: Ahn, Kwangjun, et al.
Pubblicazione: (2024)
Improved Rates for Stochastic Variance-Reduced Difference-of-Convex Algorithms
di: Nguyen, Anh Duc, et al.
Pubblicazione: (2025)
di: Nguyen, Anh Duc, et al.
Pubblicazione: (2025)
Multi-Objective LQR with Linear Scalarization
di: Jadbabaie, Ali, et al.
Pubblicazione: (2024)
di: Jadbabaie, Ali, et al.
Pubblicazione: (2024)
The Multi-Block DC Function Class: Theory, Algorithms, and Applications
di: Fatemi, Pouria, et al.
Pubblicazione: (2026)
di: Fatemi, Pouria, et al.
Pubblicazione: (2026)
Tight Generalization Bounds for Noiseless Inverse Optimization
di: Fatemi, Pouria, et al.
Pubblicazione: (2026)
di: Fatemi, Pouria, et al.
Pubblicazione: (2026)
Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification
di: Jung, Hyunji, et al.
Pubblicazione: (2025)
di: Jung, Hyunji, et al.
Pubblicazione: (2025)
Stabilizing reinforcement learning control: A modular framework for optimizing over all stable behavior
di: Lawrence, Nathan P., et al.
Pubblicazione: (2023)
di: Lawrence, Nathan P., et al.
Pubblicazione: (2023)
GLinSAT: The General Linear Satisfiability Neural Network Layer By Accelerated Gradient Descent
di: Zeng, Hongtai, et al.
Pubblicazione: (2024)
di: Zeng, Hongtai, et al.
Pubblicazione: (2024)
Randomized Block Coordinate DC Programming
di: Maskan, Hoomaan, et al.
Pubblicazione: (2024)
di: Maskan, Hoomaan, et al.
Pubblicazione: (2024)
Revisiting Frank-Wolfe for Structured Nonconvex Optimization
di: Maskan, Hoomaan, et al.
Pubblicazione: (2025)
di: Maskan, Hoomaan, et al.
Pubblicazione: (2025)
Graph Transformers Dream of Electric Flow
di: Cheng, Xiang, et al.
Pubblicazione: (2024)
di: Cheng, Xiang, et al.
Pubblicazione: (2024)
Sinkhorn doubly stochastic attention rank decay analysis
di: Lapenna, Michela, et al.
Pubblicazione: (2026)
di: Lapenna, Michela, et al.
Pubblicazione: (2026)
Provable Benefit of Random Permutations over Uniform Sampling in Stochastic Coordinate Descent
di: Kim, Donghwa, et al.
Pubblicazione: (2025)
di: Kim, Donghwa, et al.
Pubblicazione: (2025)
Availability is all you need: achieving optimal regret with minimal information for dynamic matching
di: Kerimov, Süleyman, et al.
Pubblicazione: (2025)
di: Kerimov, Süleyman, et al.
Pubblicazione: (2025)
Wasserstein gradient flow for optimal probability measure decomposition
di: Han, Jiangze, et al.
Pubblicazione: (2024)
di: Han, Jiangze, et al.
Pubblicazione: (2024)
Temporal Robustness in Discrete Time Linear Dynamical Systems
di: Metya, Nilava, et al.
Pubblicazione: (2025)
di: Metya, Nilava, et al.
Pubblicazione: (2025)
LinSATNet: The Positive Linear Satisfiability Neural Networks
di: Wang, Runzhong, et al.
Pubblicazione: (2024)
di: Wang, Runzhong, et al.
Pubblicazione: (2024)
Online Learning for Supervisory Switching Control
di: Sun, Haoyuan, et al.
Pubblicazione: (2026)
di: Sun, Haoyuan, et al.
Pubblicazione: (2026)
A least-square method for non-asymptotic identification in linear switching control
di: Sun, Haoyuan, et al.
Pubblicazione: (2024)
di: Sun, Haoyuan, et al.
Pubblicazione: (2024)
Employing Deep Neural Operators for PDE control by decoupling training and optimization
di: Lundqvist, Oliver G. S., et al.
Pubblicazione: (2025)
di: Lundqvist, Oliver G. S., et al.
Pubblicazione: (2025)
Layerwise goal-oriented adaptivity for neural ODEs: an optimal control perspective
di: Hintermüller, Michael, et al.
Pubblicazione: (2026)
di: Hintermüller, Michael, et al.
Pubblicazione: (2026)
FMIP: Joint Continuous-Integer Flow For Mixed-Integer Linear Programming
di: Li, Hongpei, et al.
Pubblicazione: (2025)
di: Li, Hongpei, et al.
Pubblicazione: (2025)
Computing Brascamp-Lieb Constants through the lens of Thompson Geometry
di: Weber, Melanie, et al.
Pubblicazione: (2022)
di: Weber, Melanie, et al.
Pubblicazione: (2022)
A deep learning method for solving stochastic optimal control problems driven by fully-coupled FBSDEs
di: Ji, Shaolin, et al.
Pubblicazione: (2022)
di: Ji, Shaolin, et al.
Pubblicazione: (2022)
Mixed variable structural optimization using mixed variable system Monte Carlo tree search formulation
di: Ko, Fu-Yao, et al.
Pubblicazione: (2023)
di: Ko, Fu-Yao, et al.
Pubblicazione: (2023)
Nesterov Acceleration with Operator Decomposition
di: Lee, Jaewook, et al.
Pubblicazione: (2026)
di: Lee, Jaewook, et al.
Pubblicazione: (2026)
Documenti analoghi
-
How to escape sharp minima with random perturbations
di: Ahn, Kwangjun, et al.
Pubblicazione: (2023) -
Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training
di: Song, Minhak, et al.
Pubblicazione: (2025) -
Does SGD really happen in tiny subspaces?
di: Song, Minhak, et al.
Pubblicazione: (2024) -
Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch Regime
di: Baek, Beomhan, et al.
Pubblicazione: (2025) -
Riemannian Bilevel Optimization
di: Dutta, Sanchayan, et al.
Pubblicazione: (2024)