Sarah Frank-Wolfe: Methods for Constrained Optimization with Best Rates and Practical Features
Fuente:
arXiv
Saved in:
| Main Authors: | Beznosikov, Aleksandr, Dobre, David, Gidel, Gauthier |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Stochastic Frank-Wolfe: Unified Analysis and Zoo of Special Cases
by: Nazykov, Ruslan, et al.
Published: (2024)
by: Nazykov, Ruslan, et al.
Published: (2024)
Ito Diffusion Approximation of Universal Ito Chains for Sampling, Optimization and Boosting
by: Ustimenko, Aleksei, et al.
Published: (2023)
by: Ustimenko, Aleksei, et al.
Published: (2023)
Where Does Warm-Up Come From? Adaptive Scheduling for Norm-Constrained Optimizers
by: Riabinin, Artem, et al.
Published: (2026)
by: Riabinin, Artem, et al.
Published: (2026)
Accelerated Methods with Compressed Communications for Distributed Optimization Problems under Data Similarity
by: Bylinkin, Dmitry, et al.
Published: (2024)
by: Bylinkin, Dmitry, et al.
Published: (2024)
Random-reshuffled SARAH does not need a full gradient computations
by: Beznosikov, Aleksandr, et al.
Published: (2021)
by: Beznosikov, Aleksandr, et al.
Published: (2021)
Lions and Muons: Optimization via Stochastic Frank-Wolfe
by: Sfyraki, Maria-Eleni, et al.
Published: (2025)
by: Sfyraki, Maria-Eleni, et al.
Published: (2025)
Beyond Short Steps in Frank-Wolfe Algorithms
by: Martínez-Rubio, David, et al.
Published: (2025)
by: Martínez-Rubio, David, et al.
Published: (2025)
Stochastic Compositional Optimization via Hybrid Momentum Frank--Wolfe
by: Chayti, El Mahdi
Published: (2026)
by: Chayti, El Mahdi
Published: (2026)
Scalable DC Optimization via Adaptive Frank-Wolfe Algorithms
by: Pokutta, Sebastian
Published: (2025)
by: Pokutta, Sebastian
Published: (2025)
Optimal Data Splitting in Distributed Optimization for Machine Learning
by: Medyakov, Daniil, et al.
Published: (2024)
by: Medyakov, Daniil, et al.
Published: (2024)
Hierarchical Mixture-of-Experts with Two-Stage Optimization
by: Molodtsov, Gleb, et al.
Published: (2026)
by: Molodtsov, Gleb, et al.
Published: (2026)
Convergence of Clipped-SGD for Convex $(L_0,L_1)$-Smooth Optimization with Heavy-Tailed Noise
by: Chezhegov, Savelii, et al.
Published: (2025)
by: Chezhegov, Savelii, et al.
Published: (2025)
Accelerated Stochastic ExtraGradient: Mixing Hessian and Gradient Similarity to Reduce Communication in Distributed and Federated Learning
by: Bylinkin, Dmitry, et al.
Published: (2024)
by: Bylinkin, Dmitry, et al.
Published: (2024)
Linear Convergence of the Frank-Wolfe Algorithm over Product Polytopes
by: Iommazzo, Gabriele, et al.
Published: (2025)
by: Iommazzo, Gabriele, et al.
Published: (2025)
Omega: Optimistic EMA Gradients
by: Ramirez, Juan, et al.
Published: (2023)
by: Ramirez, Juan, et al.
Published: (2023)
Using Taylor-Approximated Gradients to Improve the Frank-Wolfe Method for Empirical Risk Minimization
by: Xiong, Zikai, et al.
Published: (2022)
by: Xiong, Zikai, et al.
Published: (2022)
Revisiting Frank-Wolfe for Structured Nonconvex Optimization
by: Maskan, Hoomaan, et al.
Published: (2025)
by: Maskan, Hoomaan, et al.
Published: (2025)
Accelerated Frank-Wolfe Algorithms: Complementarity Conditions and Sparsity
by: Garber, Dan
Published: (2025)
by: Garber, Dan
Published: (2025)
Preconditioned Norms: A Unified Framework for Steepest Descent, Quasi-Newton and Adaptive Methods
by: Veprikov, Andrey, et al.
Published: (2025)
by: Veprikov, Andrey, et al.
Published: (2025)
Variance Reduction Methods Do Not Need to Compute Full Gradients: Improved Efficiency through Shuffling
by: Medyakov, Daniil, et al.
Published: (2025)
by: Medyakov, Daniil, et al.
Published: (2025)
When is Momentum Extragradient Optimal? A Polynomial-Based Analysis
by: Kim, Junhyung Lyle, et al.
Published: (2022)
by: Kim, Junhyung Lyle, et al.
Published: (2022)
Compression-aware Training of Neural Networks using Frank-Wolfe
by: Zimmer, Max, et al.
Published: (2022)
by: Zimmer, Max, et al.
Published: (2022)
A Randomized Linearly Convergent Frank-Wolfe-type Method for Smooth Convex Minimization over the Spectrahedron
by: Garber, Dan
Published: (2025)
by: Garber, Dan
Published: (2025)
First Order Methods with Markovian Noise: from Acceleration to Variational Inequalities
by: Beznosikov, Aleksandr, et al.
Published: (2023)
by: Beznosikov, Aleksandr, et al.
Published: (2023)
Stochastic Gradient Methods with Preconditioned Updates
by: Sadiev, Abdurakhmon, et al.
Published: (2022)
by: Sadiev, Abdurakhmon, et al.
Published: (2022)
Accelerated Methods with Complexity Separation Under Data Similarity for Federated Learning Problems
by: Bylinkin, Dmitry, et al.
Published: (2026)
by: Bylinkin, Dmitry, et al.
Published: (2026)
Dimension-adapted Momentum Outscales SGD
by: Ferbach, Damien, et al.
Published: (2025)
by: Ferbach, Damien, et al.
Published: (2025)
Logarithmic-time Schedules for Scaling Language Models with Momentum
by: Ferbach, Damien, et al.
Published: (2026)
by: Ferbach, Damien, et al.
Published: (2026)
Enhancing Stability of Physics-Informed Neural Network Training Through Saddle-Point Reformulation
by: Bylinkin, Dmitry, et al.
Published: (2025)
by: Bylinkin, Dmitry, et al.
Published: (2025)
Gradient-Free Approaches is a Key to an Efficient Interaction with Markovian Stochasticity
by: Prokhorov, Boris, et al.
Published: (2026)
by: Prokhorov, Boris, et al.
Published: (2026)
Don't Be Greedy, Just Relax! Pruning LLMs via Frank-Wolfe
by: Roux, Christophe, et al.
Published: (2025)
by: Roux, Christophe, et al.
Published: (2025)
Zero-Order Optimization for LLM Fine-Tuning via Learnable Direction Sampling
by: Parfenov, Valery, et al.
Published: (2026)
by: Parfenov, Valery, et al.
Published: (2026)
Fast Stochastic Composite Minimization and an Accelerated Frank-Wolfe Algorithm under Parallelization
by: Dubois-Taine, Benjamin, et al.
Published: (2022)
by: Dubois-Taine, Benjamin, et al.
Published: (2022)
Scalable Frank-Wolfe on Generalized Self-concordant Functions via Simple Steps
by: Carderera, Alejandro, et al.
Published: (2021)
by: Carderera, Alejandro, et al.
Published: (2021)
Projected Forward Gradient-Guided Frank-Wolfe Algorithm via Variance Reduction
by: Rostami, M., et al.
Published: (2024)
by: Rostami, M., et al.
Published: (2024)
Sign-SGD via Parameter-Free Optimization
by: Medyakov, Daniil, et al.
Published: (2025)
by: Medyakov, Daniil, et al.
Published: (2025)
Leveraging Coordinate Momentum in SignSGD and Muon: Memory-Optimized Zero-Order
by: Petrov, Egor, et al.
Published: (2025)
by: Petrov, Egor, et al.
Published: (2025)
WeightLoRA: Keep Only Necessary Adapters
by: Veprikov, Andrey, et al.
Published: (2025)
by: Veprikov, Andrey, et al.
Published: (2025)
Shuffling Heuristic in Variational Inequalities: Establishing New Convergence Guarantees
by: Medyakov, Daniil, et al.
Published: (2025)
by: Medyakov, Daniil, et al.
Published: (2025)
Solving Hidden Monotone Variational Inequalities with Surrogate Losses
by: D'Orazio, Ryan, et al.
Published: (2024)
by: D'Orazio, Ryan, et al.
Published: (2024)
Similar Items
-
Stochastic Frank-Wolfe: Unified Analysis and Zoo of Special Cases
by: Nazykov, Ruslan, et al.
Published: (2024) -
Ito Diffusion Approximation of Universal Ito Chains for Sampling, Optimization and Boosting
by: Ustimenko, Aleksei, et al.
Published: (2023) -
Where Does Warm-Up Come From? Adaptive Scheduling for Norm-Constrained Optimizers
by: Riabinin, Artem, et al.
Published: (2026) -
Accelerated Methods with Compressed Communications for Distributed Optimization Problems under Data Similarity
by: Bylinkin, Dmitry, et al.
Published: (2024) -
Random-reshuffled SARAH does not need a full gradient computations
by: Beznosikov, Aleksandr, et al.
Published: (2021)