PLUMAGE: Probabilistic Low rank Unbiased Min Variance Gradient Estimator for Efficient Large Model Training
Fuente:
arXiv
Saved in:
| Main Authors: | Haroush, Matan, Soudry, Daniel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Minimum Variance Unbiased N:M Sparsity for the Neural Gradients
by: Chmiel, Brian, et al.
Published: (2022)
by: Chmiel, Brian, et al.
Published: (2022)
OEUVRE: OnlinE Unbiased Variance-Reduced loss Estimation
by: Pardeshi, Kanad, et al.
Published: (2025)
by: Pardeshi, Kanad, et al.
Published: (2025)
Beyond ReinMax: Low-Variance Gradient Estimators for Discrete Latent Variables
by: Wang, Daniel, et al.
Published: (2026)
by: Wang, Daniel, et al.
Published: (2026)
Gradient Weight-normalized Low-rank Projection for Efficient LLM Training
by: Huang, Jia-Hong, et al.
Published: (2024)
by: Huang, Jia-Hong, et al.
Published: (2024)
Efficient and Unbiased Sampling from Boltzmann Distributions via Variance-Tuned Diffusion Models
by: Zhang, Fengzhe, et al.
Published: (2025)
by: Zhang, Fengzhe, et al.
Published: (2025)
The Implicit Bias of Gradient Descent on Separable Multiclass Data
by: Ravi, Hrithik, et al.
Published: (2024)
by: Ravi, Hrithik, et al.
Published: (2024)
Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient Estimation
by: Panferov, Andrei, et al.
Published: (2026)
by: Panferov, Andrei, et al.
Published: (2026)
Unbiased Gradient Low-Rank Projection
by: Pan, Rui, et al.
Published: (2025)
by: Pan, Rui, et al.
Published: (2025)
From Continual Learning to SGD and Back: Better Rates for Continual Linear Models
by: Evron, Itay, et al.
Published: (2025)
by: Evron, Itay, et al.
Published: (2025)
Are Greedy Task Orderings Better Than Random in Continual Linear Regression?
by: Tsipory, Matan, et al.
Published: (2025)
by: Tsipory, Matan, et al.
Published: (2025)
Low-rank Momentum Factorization for Memory Efficient Training
by: Mahdavinia, Pouria, et al.
Published: (2025)
by: Mahdavinia, Pouria, et al.
Published: (2025)
The Implicit Bias of Gradient Descent on Separable Data
by: Soudry, Daniel, et al.
Published: (2017)
by: Soudry, Daniel, et al.
Published: (2017)
Optimal Rates in Continual Linear Regression via Increasing Regularization
by: Levinstein, Ran, et al.
Published: (2025)
by: Levinstein, Ran, et al.
Published: (2025)
FP4 All the Way: Fully Quantized Training of LLMs
by: Chmiel, Brian, et al.
Published: (2025)
by: Chmiel, Brian, et al.
Published: (2025)
Gradient Estimation and Variance Reduction in Stochastic and Deterministic Models
by: Keane, Ronan
Published: (2024)
by: Keane, Ronan
Published: (2024)
Enhanced Federated Optimization: Adaptive Unbiased Client Sampling with Reduced Variance
by: Zeng, Dun, et al.
Published: (2023)
by: Zeng, Dun, et al.
Published: (2023)
Randomized Gradient Subspaces for Efficient Large Language Model Training
by: Rajabi, Sahar, et al.
Published: (2025)
by: Rajabi, Sahar, et al.
Published: (2025)
On the Optimal Construction of Unbiased Gradient Estimators for Zeroth-Order Optimization
by: Ma, Shaocong, et al.
Published: (2025)
by: Ma, Shaocong, et al.
Published: (2025)
Accurate Neural Training with 4-bit Matrix Multiplications at Standard Formats
by: Chmiel, Brian, et al.
Published: (2021)
by: Chmiel, Brian, et al.
Published: (2021)
Workspace Optimization: How to Train Your Agent
by: Sarafian, Elad, et al.
Published: (2026)
by: Sarafian, Elad, et al.
Published: (2026)
The Dimension Strikes Back with Gradients: Generalization of Gradient Methods in Stochastic Convex Optimization
by: Schliserman, Matan, et al.
Published: (2024)
by: Schliserman, Matan, et al.
Published: (2024)
Multiclass Loss Geometry Matters for Generalization of Gradient Descent in Separable Classification
by: Schliserman, Matan, et al.
Published: (2025)
by: Schliserman, Matan, et al.
Published: (2025)
Bayesian Low-rank Adaptation for Large Language Models
by: Yang, Adam X., et al.
Published: (2023)
by: Yang, Adam X., et al.
Published: (2023)
The Joint Effect of Task Similarity and Overparameterization on Catastrophic Forgetting -- An Analytical Model
by: Goldfarb, Daniel, et al.
Published: (2024)
by: Goldfarb, Daniel, et al.
Published: (2024)
Efficient Unbiased Sparsification
by: Barnes, Leighton, et al.
Published: (2024)
by: Barnes, Leighton, et al.
Published: (2024)
MARS: Unleashing the Power of Variance Reduction for Training Large Models
by: Yuan, Huizhuo, et al.
Published: (2024)
by: Yuan, Huizhuo, et al.
Published: (2024)
Unbiased Single-Queried Gradient for Combinatorial Objective
by: Sornwanee, Thanawat
Published: (2026)
by: Sornwanee, Thanawat
Published: (2026)
Exploring Variance Reduction in Importance Sampling for Efficient DNN Training
by: Kutsuna, Takuro
Published: (2025)
by: Kutsuna, Takuro
Published: (2025)
Zero-Variance Gradients for Variational Autoencoders
by: Shao, Zilei, et al.
Published: (2025)
by: Shao, Zilei, et al.
Published: (2025)
Memory-Efficient LLM Training by Various-Grained Low-Rank Projection of Gradients
by: Wang, Yezhen, et al.
Published: (2025)
by: Wang, Yezhen, et al.
Published: (2025)
GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
by: Zhao, Jiawei, et al.
Published: (2024)
by: Zhao, Jiawei, et al.
Published: (2024)
Grass: Compute Efficient Low-Memory LLM Training with Structured Sparse Gradients
by: Muhamed, Aashiq, et al.
Published: (2024)
by: Muhamed, Aashiq, et al.
Published: (2024)
Efficient and Unbiased Sampling of Boltzmann Distributions via Consistency Models
by: Zhang, Fengzhe, et al.
Published: (2024)
by: Zhang, Fengzhe, et al.
Published: (2024)
Gradient Boosted Mixed Models: Flexible Joint Estimation of Mean and Variance Components for Clustered Data
by: Prevett, Mitchell L., et al.
Published: (2025)
by: Prevett, Mitchell L., et al.
Published: (2025)
Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models
by: Zhang, Jingyang, et al.
Published: (2024)
by: Zhang, Jingyang, et al.
Published: (2024)
DR-Encoder: Encode Low-rank Gradients with Random Prior for Large Language Models Differentially Privately
by: Wu, Huiwen, et al.
Published: (2024)
by: Wu, Huiwen, et al.
Published: (2024)
Explore to Generalize in Zero-Shot RL
by: Zisselman, Ev, et al.
Published: (2023)
by: Zisselman, Ev, et al.
Published: (2023)
Foldable SuperNets: Scalable Merging of Transformers with Different Initializations and Tasks
by: Kinderman, Edan, et al.
Published: (2024)
by: Kinderman, Edan, et al.
Published: (2024)
LOST: Low-rank and Sparse Pre-training for Large Language Models
by: Li, Jiaxi, et al.
Published: (2025)
by: Li, Jiaxi, et al.
Published: (2025)
Machine Learning Predictors for Min-Entropy Estimation
by: Blanco-Romero, Javier, et al.
Published: (2024)
by: Blanco-Romero, Javier, et al.
Published: (2024)
Similar Items
-
Minimum Variance Unbiased N:M Sparsity for the Neural Gradients
by: Chmiel, Brian, et al.
Published: (2022) -
OEUVRE: OnlinE Unbiased Variance-Reduced loss Estimation
by: Pardeshi, Kanad, et al.
Published: (2025) -
Beyond ReinMax: Low-Variance Gradient Estimators for Discrete Latent Variables
by: Wang, Daniel, et al.
Published: (2026) -
Gradient Weight-normalized Low-rank Projection for Efficient LLM Training
by: Huang, Jia-Hong, et al.
Published: (2024) -
Efficient and Unbiased Sampling from Boltzmann Distributions via Variance-Tuned Diffusion Models
by: Zhang, Fengzhe, et al.
Published: (2025)