Omni-Masked Gradient Descent: Memory-Efficient Optimization via Mask Traversal with Improved Convergence
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Hui, Ren, Tao, Jiang, Jinyang, Tian, Wan, Peng, Yijie |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RiskPO: Risk-based Policy Optimization via Verifiable Reward for LLM Post-Training
by: Ren, Tao, et al.
Published: (2025)
by: Ren, Tao, et al.
Published: (2025)
Learning Provably Improves the Convergence of Gradient Descent
by: Song, Qingyu, et al.
Published: (2025)
by: Song, Qingyu, et al.
Published: (2025)
An Algebraically Converging Stochastic Gradient Descent Algorithm for Global Optimization
by: Engquist, Björn, et al.
Published: (2022)
by: Engquist, Björn, et al.
Published: (2022)
FlashMask: Efficient and Rich Mask Extension of FlashAttention
by: Wang, Guoxia, et al.
Published: (2024)
by: Wang, Guoxia, et al.
Published: (2024)
Stochastic Approximation Methods for Distortion Risk Measure Optimization
by: Jiang, Jinyang, et al.
Published: (2025)
by: Jiang, Jinyang, et al.
Published: (2025)
Optimizing Decoding Paths in Masked Diffusion Models by Quantifying Uncertainty
by: Chen, Ziyu, et al.
Published: (2025)
by: Chen, Ziyu, et al.
Published: (2025)
Increasing Batch Size Improves Convergence of Stochastic Gradient Descent with Momentum
by: Kamo, Keisuke, et al.
Published: (2025)
by: Kamo, Keisuke, et al.
Published: (2025)
Nonparametric Bayesian Optimization for General Rewards
by: Zhang, Zishi, et al.
Published: (2026)
by: Zhang, Zishi, et al.
Published: (2026)
Multi-Agent Debate with Memory Masking
by: Tian, Hongduan, et al.
Published: (2026)
by: Tian, Hongduan, et al.
Published: (2026)
Where to Mask: Structure-Guided Masking for Graph Masked Autoencoders
by: Liu, Chuang, et al.
Published: (2024)
by: Liu, Chuang, et al.
Published: (2024)
On the Convergence of Gradient Descent on Learning Transformers with Residual Connections
by: Qin, Zhen, et al.
Published: (2025)
by: Qin, Zhen, et al.
Published: (2025)
Memory-Efficient Optimization with Factorized Hamiltonian Descent
by: Nguyen, Son, et al.
Published: (2024)
by: Nguyen, Son, et al.
Published: (2024)
On the Convergence Rate of LoRA Gradient Descent
by: Mu, Siqiao, et al.
Published: (2025)
by: Mu, Siqiao, et al.
Published: (2025)
Sharp Convergence Rates for Masked Diffusion Models
by: Liang, Yuchen, et al.
Published: (2026)
by: Liang, Yuchen, et al.
Published: (2026)
Accelerating Convergence of Stein Variational Gradient Descent via Deep Unfolding
by: Kawamura, Yuya, et al.
Published: (2024)
by: Kawamura, Yuya, et al.
Published: (2024)
SeWA: Selective Weight Average via Probabilistic Masking
by: Wang, Peng, et al.
Published: (2025)
by: Wang, Peng, et al.
Published: (2025)
How to Compress KV Cache in RL Post-Training? Shadow Mask Distillation for Memory-Efficient Alignment
by: Zhu, Rui, et al.
Published: (2026)
by: Zhu, Rui, et al.
Published: (2026)
CoNNect: Connectivity-Based Regularization for Structural Pruning
by: Franssen, Christian, et al.
Published: (2025)
by: Franssen, Christian, et al.
Published: (2025)
Improved Finite-Particle Convergence Rates for Stein Variational Gradient Descent
by: Banerjee, Sayan, et al.
Published: (2024)
by: Banerjee, Sayan, et al.
Published: (2024)
Closing the Loop: Coordinating Inventory and Recommendation via Deep Reinforcement Learning on Multiple Timescales
by: Jiang, Jinyang, et al.
Published: (2025)
by: Jiang, Jinyang, et al.
Published: (2025)
Sculpting Memory: Multi-Concept Forgetting in Diffusion Models via Dynamic Mask and Concept-Aware Optimization
by: Li, Gen, et al.
Published: (2025)
by: Li, Gen, et al.
Published: (2025)
Robust Federated Learning against Noisy Clients via Masked Optimization
by: Jiang, Xuefeng, et al.
Published: (2025)
by: Jiang, Xuefeng, et al.
Published: (2025)
Accelerated Gradient Descent for Faster Convergence with Minimal Overhead
by: Graca, Manuel, et al.
Published: (2026)
by: Graca, Manuel, et al.
Published: (2026)
Machine Learning-Assisted High-Dimensional Matrix Estimation
by: Tian, Wan, et al.
Published: (2026)
by: Tian, Wan, et al.
Published: (2026)
Convergence of Alternating Gradient Descent for Matrix Factorization
by: Ward, Rachel, et al.
Published: (2023)
by: Ward, Rachel, et al.
Published: (2023)
Learning Associative Memories with Gradient Descent
by: Cabannes, Vivien, et al.
Published: (2024)
by: Cabannes, Vivien, et al.
Published: (2024)
Convergence of Spectral Descent for Non-smooth Optimization
by: Yang, Yixuan, et al.
Published: (2026)
by: Yang, Yixuan, et al.
Published: (2026)
SAU: Sparsity-Aware Unlearning for LLMs via Gradient Masking and Importance Redistribution
by: Wang, Yuze, et al.
Published: (2026)
by: Wang, Yuze, et al.
Published: (2026)
Product-Stability: Provable Convergence for Gradient Descent on the Edge of Stability
by: Gan, Eric
Published: (2026)
by: Gan, Eric
Published: (2026)
FLOPS: Forward Learning with OPtimal Sampling
by: Ren, Tao, et al.
Published: (2024)
by: Ren, Tao, et al.
Published: (2024)
Minimax Optimal Convergence of Gradient Descent in Logistic Regression via Large and Adaptive Stepsizes
by: Zhang, Ruiqi, et al.
Published: (2025)
by: Zhang, Ruiqi, et al.
Published: (2025)
On the Convergence of Gradient Descent for Large Learning Rates
by: Crăciun, Alexandru, et al.
Published: (2024)
by: Crăciun, Alexandru, et al.
Published: (2024)
Optimizing Predictive AI in Physical Design Flows with Mini Pixel Batch Gradient Descent
by: Yang, Haoyu, et al.
Published: (2024)
by: Yang, Haoyu, et al.
Published: (2024)
Beyond Masked and Unmasked: Discrete Diffusion Models via Partial Masking
by: Chao, Chen-Hao, et al.
Published: (2025)
by: Chao, Chen-Hao, et al.
Published: (2025)
Improving Robustness In Sparse Autoencoders via Masked Regularization
by: Narayanaswamy, Vivek, et al.
Published: (2026)
by: Narayanaswamy, Vivek, et al.
Published: (2026)
Gradient Routing: Masking Gradients to Localize Computation in Neural Networks
by: Cloud, Alex, et al.
Published: (2024)
by: Cloud, Alex, et al.
Published: (2024)
OmniCast: A Masked Latent Diffusion Model for Weather Forecasting Across Time Scales
by: Nguyen, Tung, et al.
Published: (2025)
by: Nguyen, Tung, et al.
Published: (2025)
Convergence Analysis of Stochastic Gradient Descent with MCMC Estimators
by: Li, Tianyou, et al.
Published: (2023)
by: Li, Tianyou, et al.
Published: (2023)
Open Problem: Anytime Convergence Rate of Gradient Descent
by: Kornowski, Guy, et al.
Published: (2024)
by: Kornowski, Guy, et al.
Published: (2024)
PolyG: Adaptive Graph Traversal for Diverse GraphRAG Questions
by: Liu, Renjie, et al.
Published: (2025)
by: Liu, Renjie, et al.
Published: (2025)
Similar Items
-
RiskPO: Risk-based Policy Optimization via Verifiable Reward for LLM Post-Training
by: Ren, Tao, et al.
Published: (2025) -
Learning Provably Improves the Convergence of Gradient Descent
by: Song, Qingyu, et al.
Published: (2025) -
An Algebraically Converging Stochastic Gradient Descent Algorithm for Global Optimization
by: Engquist, Björn, et al.
Published: (2022) -
FlashMask: Efficient and Rich Mask Extension of FlashAttention
by: Wang, Guoxia, et al.
Published: (2024) -
Stochastic Approximation Methods for Distortion Risk Measure Optimization
by: Jiang, Jinyang, et al.
Published: (2025)