WeightLoRA: Keep Only Necessary Adapters
Fuente:
arXiv
Saved in:
| Main Authors: | Veprikov, Andrey, Solodkin, Vladimir, Zyl, Alexander, Savchenko, Andrey, Beznosikov, Aleksandr |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Methods for Optimization Problems with Markovian Stochasticity and Non-Euclidean Geometry
by: Solodkin, Vladimir, et al.
Published: (2024)
by: Solodkin, Vladimir, et al.
Published: (2024)
Markovian Compression: Looking to the Past Helps Accelerate the Future
by: Veprikov, Andrey, et al.
Published: (2026)
by: Veprikov, Andrey, et al.
Published: (2026)
Where Does Warm-Up Come From? Adaptive Scheduling for Norm-Constrained Optimizers
by: Riabinin, Artem, et al.
Published: (2026)
by: Riabinin, Artem, et al.
Published: (2026)
Zero-Order Optimization for LLM Fine-Tuning via Learnable Direction Sampling
by: Parfenov, Valery, et al.
Published: (2026)
by: Parfenov, Valery, et al.
Published: (2026)
Preconditioned Norms: A Unified Framework for Steepest Descent, Quasi-Newton and Adaptive Methods
by: Veprikov, Andrey, et al.
Published: (2025)
by: Veprikov, Andrey, et al.
Published: (2025)
New Aspects of Black Box Conditional Gradient: Variance Reduction and One Point Feedback
by: Veprikov, Andrey, et al.
Published: (2024)
by: Veprikov, Andrey, et al.
Published: (2024)
Leveraging Coordinate Momentum in SignSGD and Muon: Memory-Optimized Zero-Order
by: Petrov, Egor, et al.
Published: (2025)
by: Petrov, Egor, et al.
Published: (2025)
Exploring New Frontiers in Vertical Federated Learning: the Role of Saddle Point Reformulation
by: Beznosikov, Aleksandr, et al.
Published: (2026)
by: Beznosikov, Aleksandr, et al.
Published: (2026)
Methods for Solving Variational Inequalities with Markovian Stochasticity
by: Solodkin, Vladimir, et al.
Published: (2024)
by: Solodkin, Vladimir, et al.
Published: (2024)
Accelerated Stochastic Gradient Method with Applications to Consensus Problem in Markov-Varying Networks
by: Solodkin, Vladimir, et al.
Published: (2024)
by: Solodkin, Vladimir, et al.
Published: (2024)
Stochastic Frank-Wolfe: Unified Analysis and Zoo of Special Cases
by: Nazykov, Ruslan, et al.
Published: (2024)
by: Nazykov, Ruslan, et al.
Published: (2024)
Random-reshuffled SARAH does not need a full gradient computations
by: Beznosikov, Aleksandr, et al.
Published: (2021)
by: Beznosikov, Aleksandr, et al.
Published: (2021)
Ito Diffusion Approximation of Universal Ito Chains for Sampling, Optimization and Boosting
by: Ustimenko, Aleksei, et al.
Published: (2023)
by: Ustimenko, Aleksei, et al.
Published: (2023)
DyKAF: Dynamical Kronecker Approximation of the Fisher Information Matrix for Gradient Preconditioning
by: Yudin, Nikolay, et al.
Published: (2025)
by: Yudin, Nikolay, et al.
Published: (2025)
Gradient-Free Approaches is a Key to an Efficient Interaction with Markovian Stochasticity
by: Prokhorov, Boris, et al.
Published: (2026)
by: Prokhorov, Boris, et al.
Published: (2026)
Optimal Data Splitting in Distributed Optimization for Machine Learning
by: Medyakov, Daniil, et al.
Published: (2024)
by: Medyakov, Daniil, et al.
Published: (2024)
Accelerated Stochastic ExtraGradient: Mixing Hessian and Gradient Similarity to Reduce Communication in Distributed and Federated Learning
by: Bylinkin, Dmitry, et al.
Published: (2024)
by: Bylinkin, Dmitry, et al.
Published: (2024)
Sarah Frank-Wolfe: Methods for Constrained Optimization with Best Rates and Practical Features
by: Beznosikov, Aleksandr, et al.
Published: (2023)
by: Beznosikov, Aleksandr, et al.
Published: (2023)
Hierarchical Mixture-of-Experts with Two-Stage Optimization
by: Molodtsov, Gleb, et al.
Published: (2026)
by: Molodtsov, Gleb, et al.
Published: (2026)
Extragradient Sliding for Composite Non-Monotone Variational Inequalities
by: Emelyanov, Roman, et al.
Published: (2024)
by: Emelyanov, Roman, et al.
Published: (2024)
Local SGD for Near-Quadratic Problems: Improving Convergence under Unconstrained Noise Conditions
by: Sadchikov, Andrey, et al.
Published: (2024)
by: Sadchikov, Andrey, et al.
Published: (2024)
Convergence of Clipped-SGD for Convex $(L_0,L_1)$-Smooth Optimization with Heavy-Tailed Noise
by: Chezhegov, Savelii, et al.
Published: (2025)
by: Chezhegov, Savelii, et al.
Published: (2025)
Enhancing Stability of Physics-Informed Neural Network Training Through Saddle-Point Reformulation
by: Bylinkin, Dmitry, et al.
Published: (2025)
by: Bylinkin, Dmitry, et al.
Published: (2025)
Distributed Saddle-Point Problems: Lower Bounds, Near-Optimal and Robust Algorithms
by: Beznosikov, Aleksandr, et al.
Published: (2020)
by: Beznosikov, Aleksandr, et al.
Published: (2020)
Accelerated Methods with Compressed Communications for Distributed Optimization Problems under Data Similarity
by: Bylinkin, Dmitry, et al.
Published: (2024)
by: Bylinkin, Dmitry, et al.
Published: (2024)
Beyond SGD, Without SVD: Proximal Subspace Iteration LoRA with Diagonal Fractional K-FAC
by: Almansoori, Abdulla Jasem, et al.
Published: (2026)
by: Almansoori, Abdulla Jasem, et al.
Published: (2026)
First Order Methods with Markovian Noise: from Acceleration to Variational Inequalities
by: Beznosikov, Aleksandr, et al.
Published: (2023)
by: Beznosikov, Aleksandr, et al.
Published: (2023)
Fast UCB-type algorithms for stochastic bandits with heavy and super heavy symmetric noise
by: Dorn, Yuriy, et al.
Published: (2024)
by: Dorn, Yuriy, et al.
Published: (2024)
Shuffling Heuristic in Variational Inequalities: Establishing New Convergence Guarantees
by: Medyakov, Daniil, et al.
Published: (2025)
by: Medyakov, Daniil, et al.
Published: (2025)
Variance Reduction Methods Do Not Need to Compute Full Gradients: Improved Efficiency through Shuffling
by: Medyakov, Daniil, et al.
Published: (2025)
by: Medyakov, Daniil, et al.
Published: (2025)
Activations and Gradients Compression for Model-Parallel Training
by: Rudakov, Mikhail, et al.
Published: (2024)
by: Rudakov, Mikhail, et al.
Published: (2024)
Accelerated Methods with Complexity Separation Under Data Similarity for Federated Learning Problems
by: Bylinkin, Dmitry, et al.
Published: (2026)
by: Bylinkin, Dmitry, et al.
Published: (2026)
Bant: Byzantine Antidote via Trial Function and Trust Scores
by: Molodtsov, Gleb, et al.
Published: (2025)
by: Molodtsov, Gleb, et al.
Published: (2025)
Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-Tailed
by: Chezhegov, Savelii, et al.
Published: (2024)
by: Chezhegov, Savelii, et al.
Published: (2024)
Stochastic Gradient Methods with Preconditioned Updates
by: Sadiev, Abdurakhmon, et al.
Published: (2022)
by: Sadiev, Abdurakhmon, et al.
Published: (2022)
A Mathematical Model of the Hidden Feedback Loop Effect in Machine Learning Systems
by: Veprikov, Andrey, et al.
Published: (2024)
by: Veprikov, Andrey, et al.
Published: (2024)
Sign-SGD via Parameter-Free Optimization
by: Medyakov, Daniil, et al.
Published: (2025)
by: Medyakov, Daniil, et al.
Published: (2025)
Similarity, Compression and Local Steps: Three Pillars of Efficient Communications for Distributed Variational Inequalities
by: Beznosikov, Aleksandr, et al.
Published: (2023)
by: Beznosikov, Aleksandr, et al.
Published: (2023)
LoRA Training in the NTK Regime has No Spurious Local Minima
by: Jang, Uijeong, et al.
Published: (2024)
by: Jang, Uijeong, et al.
Published: (2024)
Faster Than SVD, Smarter Than SGD: The OPLoRA Alternating Update
by: Almansoori, Abdulla Jasem, et al.
Published: (2025)
by: Almansoori, Abdulla Jasem, et al.
Published: (2025)
Similar Items
-
Methods for Optimization Problems with Markovian Stochasticity and Non-Euclidean Geometry
by: Solodkin, Vladimir, et al.
Published: (2024) -
Markovian Compression: Looking to the Past Helps Accelerate the Future
by: Veprikov, Andrey, et al.
Published: (2026) -
Where Does Warm-Up Come From? Adaptive Scheduling for Norm-Constrained Optimizers
by: Riabinin, Artem, et al.
Published: (2026) -
Zero-Order Optimization for LLM Fine-Tuning via Learnable Direction Sampling
by: Parfenov, Valery, et al.
Published: (2026) -
Preconditioned Norms: A Unified Framework for Steepest Descent, Quasi-Newton and Adaptive Methods
by: Veprikov, Andrey, et al.
Published: (2025)