Toward generalizable learning of all (linear) first-order methods via memory augmented Transformers
Fuente:
arXiv
Salvato in:
| Autori principali: | Dutta, Sanchayan, Sra, Suvrit |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Riemannian Bilevel Optimization
di: Dutta, Sanchayan, et al.
Pubblicazione: (2024)
di: Dutta, Sanchayan, et al.
Pubblicazione: (2024)
Linearly Convergent Algorithms for Nonsmooth Problems with Unknown Smooth Pieces
di: Zhang, Zhe, et al.
Pubblicazione: (2025)
di: Zhang, Zhe, et al.
Pubblicazione: (2025)
Implicit Bias in Matrix Factorization and its Explicit Realization in a New Architecture
di: Hou, Yikun, et al.
Pubblicazione: (2025)
di: Hou, Yikun, et al.
Pubblicazione: (2025)
How to escape sharp minima with random perturbations
di: Ahn, Kwangjun, et al.
Pubblicazione: (2023)
di: Ahn, Kwangjun, et al.
Pubblicazione: (2023)
Cost-Driven Representation Learning for Linear Quadratic Gaussian Control: Part I
di: Tian, Yi, et al.
Pubblicazione: (2022)
di: Tian, Yi, et al.
Pubblicazione: (2022)
Cost-Driven Representation Learning for Linear Quadratic Gaussian Control: Part II
di: Tian, Yi, et al.
Pubblicazione: (2026)
di: Tian, Yi, et al.
Pubblicazione: (2026)
Tight Generalization Bounds for Noiseless Inverse Optimization
di: Fatemi, Pouria, et al.
Pubblicazione: (2026)
di: Fatemi, Pouria, et al.
Pubblicazione: (2026)
Linear attention is (maybe) all you need (to understand transformer optimization)
di: Ahn, Kwangjun, et al.
Pubblicazione: (2023)
di: Ahn, Kwangjun, et al.
Pubblicazione: (2023)
First-Order Methods for Linearly Constrained Bilevel Optimization
di: Kornowski, Guy, et al.
Pubblicazione: (2024)
di: Kornowski, Guy, et al.
Pubblicazione: (2024)
Revisiting Frank-Wolfe for Structured Nonconvex Optimization
di: Maskan, Hoomaan, et al.
Pubblicazione: (2025)
di: Maskan, Hoomaan, et al.
Pubblicazione: (2025)
An adaptively inexact first-order method for bilevel optimization with application to hyperparameter learning
di: Salehi, Mohammad Sadegh, et al.
Pubblicazione: (2023)
di: Salehi, Mohammad Sadegh, et al.
Pubblicazione: (2023)
Stochastic first-order methods for average-reward Markov decision processes
di: Li, Tianjiao, et al.
Pubblicazione: (2022)
di: Li, Tianjiao, et al.
Pubblicazione: (2022)
Data augmentation for machine learning of chemical process flowsheets
di: Balhorn, Lukas Schulze, et al.
Pubblicazione: (2023)
di: Balhorn, Lukas Schulze, et al.
Pubblicazione: (2023)
The inexact power augmented Lagrangian method for constrained nonconvex optimization
di: Bodard, Alexander, et al.
Pubblicazione: (2024)
di: Bodard, Alexander, et al.
Pubblicazione: (2024)
Approximate non-linear model predictive control with safety-augmented neural networks
di: Hose, Henrik, et al.
Pubblicazione: (2023)
di: Hose, Henrik, et al.
Pubblicazione: (2023)
Training robust and generalizable quantum models
di: Berberich, Julian, et al.
Pubblicazione: (2023)
di: Berberich, Julian, et al.
Pubblicazione: (2023)
Conditionally adaptive augmented Lagrangian method for physics-informed learning of forward and inverse problems
di: Hu, Qifeng, et al.
Pubblicazione: (2025)
di: Hu, Qifeng, et al.
Pubblicazione: (2025)
Robust stochastic first order methods in heavy-tailed noise via medoid mini-batch gradient sampling
di: Vukovic, Manojlo, et al.
Pubblicazione: (2026)
di: Vukovic, Manojlo, et al.
Pubblicazione: (2026)
Improved Rates for Stochastic Variance-Reduced Difference-of-Convex Algorithms
di: Nguyen, Anh Duc, et al.
Pubblicazione: (2025)
di: Nguyen, Anh Duc, et al.
Pubblicazione: (2025)
Local linear convergence of gradient methods for overparameterized Gaussian mixtures
di: Wang, Jingxing, et al.
Pubblicazione: (2026)
di: Wang, Jingxing, et al.
Pubblicazione: (2026)
State evolution beyond first-order methods I: Rigorous predictions and finite-sample guarantees
di: Celentano, Michael, et al.
Pubblicazione: (2025)
di: Celentano, Michael, et al.
Pubblicazione: (2025)
Tight analyses of first-order methods with error feedback
di: Thomsen, Daniel Berg, et al.
Pubblicazione: (2025)
di: Thomsen, Daniel Berg, et al.
Pubblicazione: (2025)
The Multi-Block DC Function Class: Theory, Algorithms, and Applications
di: Fatemi, Pouria, et al.
Pubblicazione: (2026)
di: Fatemi, Pouria, et al.
Pubblicazione: (2026)
An accelerated first-order regularized momentum descent ascent algorithm for stochastic nonconvex-concave minimax problems
di: Zhang, Huiling, et al.
Pubblicazione: (2023)
di: Zhang, Huiling, et al.
Pubblicazione: (2023)
First-order methods for Stochastic Variational Inequality problems with Function Constraints
di: Boob, Digvijay, et al.
Pubblicazione: (2023)
di: Boob, Digvijay, et al.
Pubblicazione: (2023)
A least-square method for non-asymptotic identification in linear switching control
di: Sun, Haoyuan, et al.
Pubblicazione: (2024)
di: Sun, Haoyuan, et al.
Pubblicazione: (2024)
A distributed semismooth Newton based augmented Lagrangian method for distributed optimization
di: Ma, Qihao, et al.
Pubblicazione: (2026)
di: Ma, Qihao, et al.
Pubblicazione: (2026)
A short proof of near-linear convergence of adaptive gradient descent under fourth-order growth and convexity
di: Davis, Damek, et al.
Pubblicazione: (2026)
di: Davis, Damek, et al.
Pubblicazione: (2026)
Reinforcement learning for adaptive interior point methods in convex quadratic programming
di: Bertoncini, Jeremy, et al.
Pubblicazione: (2025)
di: Bertoncini, Jeremy, et al.
Pubblicazione: (2025)
A Minimax-MDP Framework with Future-imposed Conditions for Learning-augmented Problems
di: Chen, Xin, et al.
Pubblicazione: (2025)
di: Chen, Xin, et al.
Pubblicazione: (2025)
When GNNs meet symmetry in ILPs: an orbit-based feature augmentation approach
di: Chen, Qian, et al.
Pubblicazione: (2025)
di: Chen, Qian, et al.
Pubblicazione: (2025)
Hyperparameter tuning via trajectory predictions: Stochastic prox-linear methods in matrix sensing
di: Lou, Mengqi, et al.
Pubblicazione: (2024)
di: Lou, Mengqi, et al.
Pubblicazione: (2024)
Joint learning of a network of linear dynamical systems via total variation penalization
di: Donnat, Claire, et al.
Pubblicazione: (2025)
di: Donnat, Claire, et al.
Pubblicazione: (2025)
Randomized Block Coordinate DC Programming
di: Maskan, Hoomaan, et al.
Pubblicazione: (2024)
di: Maskan, Hoomaan, et al.
Pubblicazione: (2024)
Linear quadratic control of nonlinear systems with Koopman operator learning and the Nyström method
di: Caldarelli, Edoardo, et al.
Pubblicazione: (2024)
di: Caldarelli, Edoardo, et al.
Pubblicazione: (2024)
SGD with memory: fundamental properties and stochastic acceleration
di: Yarotsky, Dmitry, et al.
Pubblicazione: (2024)
di: Yarotsky, Dmitry, et al.
Pubblicazione: (2024)
Towards Exact Gradient-based Training on Analog In-memory Computing
di: Wu, Zhaoxian, et al.
Pubblicazione: (2024)
di: Wu, Zhaoxian, et al.
Pubblicazione: (2024)
Adaptive multi-gradient methods for quasiconvex vector optimization and applications to multi-task learning
di: Minh, Nguyen Anh, et al.
Pubblicazione: (2024)
di: Minh, Nguyen Anh, et al.
Pubblicazione: (2024)
PEPit: computer-assisted worst-case analyses of first-order optimization methods in Python
di: Goujaud, Baptiste, et al.
Pubblicazione: (2022)
di: Goujaud, Baptiste, et al.
Pubblicazione: (2022)
A successive approximation method in functional spaces for hierarchical optimal control problems and its application to learning
di: Befekadu, Getachew K.
Pubblicazione: (2024)
di: Befekadu, Getachew K.
Pubblicazione: (2024)
Documenti analoghi
-
Riemannian Bilevel Optimization
di: Dutta, Sanchayan, et al.
Pubblicazione: (2024) -
Linearly Convergent Algorithms for Nonsmooth Problems with Unknown Smooth Pieces
di: Zhang, Zhe, et al.
Pubblicazione: (2025) -
Implicit Bias in Matrix Factorization and its Explicit Realization in a New Architecture
di: Hou, Yikun, et al.
Pubblicazione: (2025) -
How to escape sharp minima with random perturbations
di: Ahn, Kwangjun, et al.
Pubblicazione: (2023) -
Cost-Driven Representation Learning for Linear Quadratic Gaussian Control: Part I
di: Tian, Yi, et al.
Pubblicazione: (2022)