Faster Than SVD, Smarter Than SGD: The OPLoRA Alternating Update
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Almansoori, Abdulla Jasem, Ivanova, Maria, Veprikov, Andrey, Beznosikov, Aleksandr, Horváth, Samuel, Takáč, Martin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond SGD, Without SVD: Proximal Subspace Iteration LoRA with Diagonal Fractional K-FAC
von: Almansoori, Abdulla Jasem, et al.
Veröffentlicht: (2026)
von: Almansoori, Abdulla Jasem, et al.
Veröffentlicht: (2026)
Collaborative and Efficient Personalization with Mixtures of Adaptors
von: Almansoori, Abdulla Jasem, et al.
Veröffentlicht: (2024)
von: Almansoori, Abdulla Jasem, et al.
Veröffentlicht: (2024)
PaDPaF: Partial Disentanglement with Partially-Federated GANs
von: Almansoori, Abdulla Jasem, et al.
Veröffentlicht: (2022)
von: Almansoori, Abdulla Jasem, et al.
Veröffentlicht: (2022)
Stochastic Gradient Methods with Preconditioned Updates
von: Sadiev, Abdurakhmon, et al.
Veröffentlicht: (2022)
von: Sadiev, Abdurakhmon, et al.
Veröffentlicht: (2022)
LionMuon: Alternating Spectral and Sign Descent for Efficient Training
von: Bolatov, Arman, et al.
Veröffentlicht: (2026)
von: Bolatov, Arman, et al.
Veröffentlicht: (2026)
Preconditioned Norms: A Unified Framework for Steepest Descent, Quasi-Newton and Adaptive Methods
von: Veprikov, Andrey, et al.
Veröffentlicht: (2025)
von: Veprikov, Andrey, et al.
Veröffentlicht: (2025)
Where Does Warm-Up Come From? Adaptive Scheduling for Norm-Constrained Optimizers
von: Riabinin, Artem, et al.
Veröffentlicht: (2026)
von: Riabinin, Artem, et al.
Veröffentlicht: (2026)
WeightLoRA: Keep Only Necessary Adapters
von: Veprikov, Andrey, et al.
Veröffentlicht: (2025)
von: Veprikov, Andrey, et al.
Veröffentlicht: (2025)
FRUGAL: Memory-Efficient Optimization by Reducing State Overhead for Scalable Training
von: Zmushko, Philip, et al.
Veröffentlicht: (2024)
von: Zmushko, Philip, et al.
Veröffentlicht: (2024)
Random-reshuffled SARAH does not need a full gradient computations
von: Beznosikov, Aleksandr, et al.
Veröffentlicht: (2021)
von: Beznosikov, Aleksandr, et al.
Veröffentlicht: (2021)
Convergence of Clipped-SGD for Convex $(L_0,L_1)$-Smooth Optimization with Heavy-Tailed Noise
von: Chezhegov, Savelii, et al.
Veröffentlicht: (2025)
von: Chezhegov, Savelii, et al.
Veröffentlicht: (2025)
Leveraging Coordinate Momentum in SignSGD and Muon: Memory-Optimized Zero-Order
von: Petrov, Egor, et al.
Veröffentlicht: (2025)
von: Petrov, Egor, et al.
Veröffentlicht: (2025)
Softsign: Smooth Sign in Your Optimizer For Better Parameter Heterogeneity Handling
von: Feoktistov, Dmitrii, et al.
Veröffentlicht: (2026)
von: Feoktistov, Dmitrii, et al.
Veröffentlicht: (2026)
Revisiting LocalSGD and SCAFFOLD: Improved Rates and Missing Analysis
von: Luo, Ruichen, et al.
Veröffentlicht: (2025)
von: Luo, Ruichen, et al.
Veröffentlicht: (2025)
Similarity, Compression and Local Steps: Three Pillars of Efficient Communications for Distributed Variational Inequalities
von: Beznosikov, Aleksandr, et al.
Veröffentlicht: (2023)
von: Beznosikov, Aleksandr, et al.
Veröffentlicht: (2023)
Zero-Order Optimization for LLM Fine-Tuning via Learnable Direction Sampling
von: Parfenov, Valery, et al.
Veröffentlicht: (2026)
von: Parfenov, Valery, et al.
Veröffentlicht: (2026)
Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-Tailed
von: Chezhegov, Savelii, et al.
Veröffentlicht: (2024)
von: Chezhegov, Savelii, et al.
Veröffentlicht: (2024)
Aligning Distributionally Robust Optimization with Practical Deep Learning Needs
von: Feoktistov, Dmitrii, et al.
Veröffentlicht: (2025)
von: Feoktistov, Dmitrii, et al.
Veröffentlicht: (2025)
OPLoRA: Orthogonal Projection LoRA Prevents Catastrophic Forgetting during Parameter-Efficient Fine-Tuning
von: Xiong, Yifeng, et al.
Veröffentlicht: (2025)
von: Xiong, Yifeng, et al.
Veröffentlicht: (2025)
On Biased Compression for Distributed Learning
von: Beznosikov, Aleksandr, et al.
Veröffentlicht: (2020)
von: Beznosikov, Aleksandr, et al.
Veröffentlicht: (2020)
Benchmarks Saturate When The Model Gets Smarter Than The Judge
von: Ballon, Marthe, et al.
Veröffentlicht: (2026)
von: Ballon, Marthe, et al.
Veröffentlicht: (2026)
Run LoRA Run: Faster and Lighter LoRA Implementations
von: Cherniuk, Daria, et al.
Veröffentlicht: (2023)
von: Cherniuk, Daria, et al.
Veröffentlicht: (2023)
Reasoning with Sampling: Your Base Model is Smarter Than You Think
von: Karan, Aayush, et al.
Veröffentlicht: (2025)
von: Karan, Aayush, et al.
Veröffentlicht: (2025)
Generalising Battery Control in Net-Zero Buildings via Personalised Federated RL
von: Avila, Nicolas M Cuadrado, et al.
Veröffentlicht: (2024)
von: Avila, Nicolas M Cuadrado, et al.
Veröffentlicht: (2024)
Byzantine-Robust Optimization under $(L_0, L_1)$-Smoothness
von: Bolatov, Arman, et al.
Veröffentlicht: (2026)
von: Bolatov, Arman, et al.
Veröffentlicht: (2026)
Thinking like a CHEMIST: Combined Heterogeneous Embedding Model Integrating Structure and Tokens
von: Rekut, Nikolai, et al.
Veröffentlicht: (2025)
von: Rekut, Nikolai, et al.
Veröffentlicht: (2025)
New Aspects of Black Box Conditional Gradient: Variance Reduction and One Point Feedback
von: Veprikov, Andrey, et al.
Veröffentlicht: (2024)
von: Veprikov, Andrey, et al.
Veröffentlicht: (2024)
LoRA Is Slower Than You Think
von: Ko, Seokmin
Veröffentlicht: (2025)
von: Ko, Seokmin
Veröffentlicht: (2025)
Sign-SGD via Parameter-Free Optimization
von: Medyakov, Daniil, et al.
Veröffentlicht: (2025)
von: Medyakov, Daniil, et al.
Veröffentlicht: (2025)
Federated Learning Can Find Friends That Are Advantageous
von: Tupitsa, Nazarii, et al.
Veröffentlicht: (2024)
von: Tupitsa, Nazarii, et al.
Veröffentlicht: (2024)
Generalized Policy Learning for Smart Grids: FL TRPO Approach
von: Li, Yunxiang, et al.
Veröffentlicht: (2024)
von: Li, Yunxiang, et al.
Veröffentlicht: (2024)
DyKAF: Dynamical Kronecker Approximation of the Fisher Information Matrix for Gradient Preconditioning
von: Yudin, Nikolay, et al.
Veröffentlicht: (2025)
von: Yudin, Nikolay, et al.
Veröffentlicht: (2025)
FedPeWS: Personalized Warmup via Subnetworks for Enhanced Heterogeneous Federated Learning
von: Tastan, Nurbek, et al.
Veröffentlicht: (2024)
von: Tastan, Nurbek, et al.
Veröffentlicht: (2024)
FERRET: Private Deep Learning Faster And Better Than DPSGD
von: Zagardo, David
Veröffentlicht: (2025)
von: Zagardo, David
Veröffentlicht: (2025)
Ito Diffusion Approximation of Universal Ito Chains for Sampling, Optimization and Boosting
von: Ustimenko, Aleksei, et al.
Veröffentlicht: (2023)
von: Ustimenko, Aleksei, et al.
Veröffentlicht: (2023)
Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints
von: Bowyer, Sam, et al.
Veröffentlicht: (2025)
von: Bowyer, Sam, et al.
Veröffentlicht: (2025)
LoFT: Low-Rank Adaptation That Behaves Like Full Fine-Tuning
von: Tastan, Nurbek, et al.
Veröffentlicht: (2025)
von: Tastan, Nurbek, et al.
Veröffentlicht: (2025)
Scalable Knowledge Editing for Mixture-of-Experts LLMs via Tensor-Structured Updates
von: Maksimov, Roman, et al.
Veröffentlicht: (2026)
von: Maksimov, Roman, et al.
Veröffentlicht: (2026)
Solving Dense Linear Systems Faster Than via Preconditioning
von: Dereziński, Michał, et al.
Veröffentlicht: (2023)
von: Dereziński, Michał, et al.
Veröffentlicht: (2023)
Exploring New Frontiers in Vertical Federated Learning: the Role of Saddle Point Reformulation
von: Beznosikov, Aleksandr, et al.
Veröffentlicht: (2026)
von: Beznosikov, Aleksandr, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Beyond SGD, Without SVD: Proximal Subspace Iteration LoRA with Diagonal Fractional K-FAC
von: Almansoori, Abdulla Jasem, et al.
Veröffentlicht: (2026) -
Collaborative and Efficient Personalization with Mixtures of Adaptors
von: Almansoori, Abdulla Jasem, et al.
Veröffentlicht: (2024) -
PaDPaF: Partial Disentanglement with Partially-Federated GANs
von: Almansoori, Abdulla Jasem, et al.
Veröffentlicht: (2022) -
Stochastic Gradient Methods with Preconditioned Updates
von: Sadiev, Abdurakhmon, et al.
Veröffentlicht: (2022) -
LionMuon: Alternating Spectral and Sign Descent for Efficient Training
von: Bolatov, Arman, et al.
Veröffentlicht: (2026)