Softsign: Smooth Sign in Your Optimizer For Better Parameter Heterogeneity Handling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Feoktistov, Dmitrii, Belinsky, Timofey, Veprikov, Andrey, Zainullin, Amir, Beznosikov, Aleksandr |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Aligning Distributionally Robust Optimization with Practical Deep Learning Needs
von: Feoktistov, Dmitrii, et al.
Veröffentlicht: (2025)
von: Feoktistov, Dmitrii, et al.
Veröffentlicht: (2025)
Leveraging Coordinate Momentum in SignSGD and Muon: Memory-Optimized Zero-Order
von: Petrov, Egor, et al.
Veröffentlicht: (2025)
von: Petrov, Egor, et al.
Veröffentlicht: (2025)
Where Does Warm-Up Come From? Adaptive Scheduling for Norm-Constrained Optimizers
von: Riabinin, Artem, et al.
Veröffentlicht: (2026)
von: Riabinin, Artem, et al.
Veröffentlicht: (2026)
LionMuon: Alternating Spectral and Sign Descent for Efficient Training
von: Bolatov, Arman, et al.
Veröffentlicht: (2026)
von: Bolatov, Arman, et al.
Veröffentlicht: (2026)
WeightLoRA: Keep Only Necessary Adapters
von: Veprikov, Andrey, et al.
Veröffentlicht: (2025)
von: Veprikov, Andrey, et al.
Veröffentlicht: (2025)
Zero-Order Optimization for LLM Fine-Tuning via Learnable Direction Sampling
von: Parfenov, Valery, et al.
Veröffentlicht: (2026)
von: Parfenov, Valery, et al.
Veröffentlicht: (2026)
Lower Bounds and Optimal Algorithms for Non-Smooth Convex Decentralized Optimization over Time-Varying Networks
von: Kovalev, Dmitry, et al.
Veröffentlicht: (2024)
von: Kovalev, Dmitry, et al.
Veröffentlicht: (2024)
Preconditioned Norms: A Unified Framework for Steepest Descent, Quasi-Newton and Adaptive Methods
von: Veprikov, Andrey, et al.
Veröffentlicht: (2025)
von: Veprikov, Andrey, et al.
Veröffentlicht: (2025)
Beyond SGD, Without SVD: Proximal Subspace Iteration LoRA with Diagonal Fractional K-FAC
von: Almansoori, Abdulla Jasem, et al.
Veröffentlicht: (2026)
von: Almansoori, Abdulla Jasem, et al.
Veröffentlicht: (2026)
Faster Than SVD, Smarter Than SGD: The OPLoRA Alternating Update
von: Almansoori, Abdulla Jasem, et al.
Veröffentlicht: (2025)
von: Almansoori, Abdulla Jasem, et al.
Veröffentlicht: (2025)
Sign-SGD via Parameter-Free Optimization
von: Medyakov, Daniil, et al.
Veröffentlicht: (2025)
von: Medyakov, Daniil, et al.
Veröffentlicht: (2025)
Benchmarking Optimizers for MLPs in Tabular Deep Learning
von: Gorishniy, Yury, et al.
Veröffentlicht: (2026)
von: Gorishniy, Yury, et al.
Veröffentlicht: (2026)
Convergence of Clipped-SGD for Convex $(L_0,L_1)$-Smooth Optimization with Heavy-Tailed Noise
von: Chezhegov, Savelii, et al.
Veröffentlicht: (2025)
von: Chezhegov, Savelii, et al.
Veröffentlicht: (2025)
Ito Diffusion Approximation of Universal Ito Chains for Sampling, Optimization and Boosting
von: Ustimenko, Aleksei, et al.
Veröffentlicht: (2023)
von: Ustimenko, Aleksei, et al.
Veröffentlicht: (2023)
New Aspects of Black Box Conditional Gradient: Variance Reduction and One Point Feedback
von: Veprikov, Andrey, et al.
Veröffentlicht: (2024)
von: Veprikov, Andrey, et al.
Veröffentlicht: (2024)
Methods for Optimization Problems with Markovian Stochasticity and Non-Euclidean Geometry
von: Solodkin, Vladimir, et al.
Veröffentlicht: (2024)
von: Solodkin, Vladimir, et al.
Veröffentlicht: (2024)
Accelerated Methods with Compressed Communications for Distributed Optimization Problems under Data Similarity
von: Bylinkin, Dmitry, et al.
Veröffentlicht: (2024)
von: Bylinkin, Dmitry, et al.
Veröffentlicht: (2024)
DyKAF: Dynamical Kronecker Approximation of the Fisher Information Matrix for Gradient Preconditioning
von: Yudin, Nikolay, et al.
Veröffentlicht: (2025)
von: Yudin, Nikolay, et al.
Veröffentlicht: (2025)
Sarah Frank-Wolfe: Methods for Constrained Optimization with Best Rates and Practical Features
von: Beznosikov, Aleksandr, et al.
Veröffentlicht: (2023)
von: Beznosikov, Aleksandr, et al.
Veröffentlicht: (2023)
Random-reshuffled SARAH does not need a full gradient computations
von: Beznosikov, Aleksandr, et al.
Veröffentlicht: (2021)
von: Beznosikov, Aleksandr, et al.
Veröffentlicht: (2021)
Hierarchical Mixture-of-Experts with Two-Stage Optimization
von: Molodtsov, Gleb, et al.
Veröffentlicht: (2026)
von: Molodtsov, Gleb, et al.
Veröffentlicht: (2026)
FRUGAL: Memory-Efficient Optimization by Reducing State Overhead for Scalable Training
von: Zmushko, Philip, et al.
Veröffentlicht: (2024)
von: Zmushko, Philip, et al.
Veröffentlicht: (2024)
A Mathematical Model of the Hidden Feedback Loop Effect in Machine Learning Systems
von: Veprikov, Andrey, et al.
Veröffentlicht: (2024)
von: Veprikov, Andrey, et al.
Veröffentlicht: (2024)
Sign Operator for Coping with Heavy-Tailed Noise in Non-Convex Optimization: High Probability Bounds Under $(L_0, L_1)$-Smoothness
von: Kornilov, Nikita, et al.
Veröffentlicht: (2025)
von: Kornilov, Nikita, et al.
Veröffentlicht: (2025)
Optimal Data Splitting in Distributed Optimization for Machine Learning
von: Medyakov, Daniil, et al.
Veröffentlicht: (2024)
von: Medyakov, Daniil, et al.
Veröffentlicht: (2024)
HARP: Hadamard-Preconditioned Adaptive Rotation Processor for Extreme LLM Quantization
von: Zagitov, Artur, et al.
Veröffentlicht: (2026)
von: Zagitov, Artur, et al.
Veröffentlicht: (2026)
Accelerated Stochastic ExtraGradient: Mixing Hessian and Gradient Similarity to Reduce Communication in Distributed and Federated Learning
von: Bylinkin, Dmitry, et al.
Veröffentlicht: (2024)
von: Bylinkin, Dmitry, et al.
Veröffentlicht: (2024)
Label Privacy in Split Learning for Large Models with Parameter-Efficient Training
von: Zmushko, Philip, et al.
Veröffentlicht: (2024)
von: Zmushko, Philip, et al.
Veröffentlicht: (2024)
Thinking like a CHEMIST: Combined Heterogeneous Embedding Model Integrating Structure and Tokens
von: Rekut, Nikolai, et al.
Veröffentlicht: (2025)
von: Rekut, Nikolai, et al.
Veröffentlicht: (2025)
Markovian Compression: Looking to the Past Helps Accelerate the Future
von: Veprikov, Andrey, et al.
Veröffentlicht: (2026)
von: Veprikov, Andrey, et al.
Veröffentlicht: (2026)
Distributed Saddle-Point Problems: Lower Bounds, Near-Optimal and Robust Algorithms
von: Beznosikov, Aleksandr, et al.
Veröffentlicht: (2020)
von: Beznosikov, Aleksandr, et al.
Veröffentlicht: (2020)
Gradient-Free Approaches is a Key to an Efficient Interaction with Markovian Stochasticity
von: Prokhorov, Boris, et al.
Veröffentlicht: (2026)
von: Prokhorov, Boris, et al.
Veröffentlicht: (2026)
Enhancing Stability of Physics-Informed Neural Network Training Through Saddle-Point Reformulation
von: Bylinkin, Dmitry, et al.
Veröffentlicht: (2025)
von: Bylinkin, Dmitry, et al.
Veröffentlicht: (2025)
Model-based Policy Optimization using Symbolic World Model
von: Gorodetskiy, Andrey, et al.
Veröffentlicht: (2024)
von: Gorodetskiy, Andrey, et al.
Veröffentlicht: (2024)
Similarity, Compression and Local Steps: Three Pillars of Efficient Communications for Distributed Variational Inequalities
von: Beznosikov, Aleksandr, et al.
Veröffentlicht: (2023)
von: Beznosikov, Aleksandr, et al.
Veröffentlicht: (2023)
Safe Continual Reinforcement Learning under Nonstationarity via Adaptive Safety Constraints
von: Tomashevskiy, Timofey
Veröffentlicht: (2026)
von: Tomashevskiy, Timofey
Veröffentlicht: (2026)
Learning to Handle Parameter Perturbations in Combinatorial Optimization: an Application to Facility Location
von: Lodi, Andrea, et al.
Veröffentlicht: (2019)
von: Lodi, Andrea, et al.
Veröffentlicht: (2019)
Shuffling Heuristic in Variational Inequalities: Establishing New Convergence Guarantees
von: Medyakov, Daniil, et al.
Veröffentlicht: (2025)
von: Medyakov, Daniil, et al.
Veröffentlicht: (2025)
Variance Reduction Methods Do Not Need to Compute Full Gradients: Improved Efficiency through Shuffling
von: Medyakov, Daniil, et al.
Veröffentlicht: (2025)
von: Medyakov, Daniil, et al.
Veröffentlicht: (2025)
Activations and Gradients Compression for Model-Parallel Training
von: Rudakov, Mikhail, et al.
Veröffentlicht: (2024)
von: Rudakov, Mikhail, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Aligning Distributionally Robust Optimization with Practical Deep Learning Needs
von: Feoktistov, Dmitrii, et al.
Veröffentlicht: (2025) -
Leveraging Coordinate Momentum in SignSGD and Muon: Memory-Optimized Zero-Order
von: Petrov, Egor, et al.
Veröffentlicht: (2025) -
Where Does Warm-Up Come From? Adaptive Scheduling for Norm-Constrained Optimizers
von: Riabinin, Artem, et al.
Veröffentlicht: (2026) -
LionMuon: Alternating Spectral and Sign Descent for Efficient Training
von: Bolatov, Arman, et al.
Veröffentlicht: (2026) -
WeightLoRA: Keep Only Necessary Adapters
von: Veprikov, Andrey, et al.
Veröffentlicht: (2025)