Minimum Variance Unbiased N:M Sparsity for the Neural Gradients
Fuente:
arXiv
Saved in:
| Main Authors: | Chmiel, Brian, Hubara, Itay, Banner, Ron, Soudry, Daniel |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FP4 All the Way: Fully Quantized Training of LLMs
by: Chmiel, Brian, et al.
Published: (2025)
by: Chmiel, Brian, et al.
Published: (2025)
Scaling FP8 training to trillion-token LLMs
by: Fishman, Maxim, et al.
Published: (2024)
by: Fishman, Maxim, et al.
Published: (2024)
Normalized Architectures are Natively 4-Bit
by: Fishman, Maxim, et al.
Published: (2026)
by: Fishman, Maxim, et al.
Published: (2026)
Towards Cheaper Inference in Deep Networks with Lower Bit-Width Accumulators
by: Blumenfeld, Yaniv, et al.
Published: (2024)
by: Blumenfeld, Yaniv, et al.
Published: (2024)
Accurate Neural Training with 4-bit Matrix Multiplications at Standard Formats
by: Chmiel, Brian, et al.
Published: (2021)
by: Chmiel, Brian, et al.
Published: (2021)
Workspace Optimization: How to Train Your Agent
by: Sarafian, Elad, et al.
Published: (2026)
by: Sarafian, Elad, et al.
Published: (2026)
Toward Efficient Permutation for Hierarchical N:M Sparsity on GPUs
by: Yu, Seungmin, et al.
Published: (2024)
by: Yu, Seungmin, et al.
Published: (2024)
EXAQ: Exponent Aware Quantization For LLMs Acceleration
by: Shkolnik, Moran, et al.
Published: (2024)
by: Shkolnik, Moran, et al.
Published: (2024)
Foldable SuperNets: Scalable Merging of Transformers with Different Initializations and Tasks
by: Kinderman, Edan, et al.
Published: (2024)
by: Kinderman, Edan, et al.
Published: (2024)
Amber Pruner: Leveraging N:M Activation Sparsity for Efficient Prefill in Large Language Models
by: An, Tai, et al.
Published: (2025)
by: An, Tai, et al.
Published: (2025)
PLUMAGE: Probabilistic Low rank Unbiased Min Variance Gradient Estimator for Efficient Large Model Training
by: Haroush, Matan, et al.
Published: (2025)
by: Haroush, Matan, et al.
Published: (2025)
E-Sparse: Boosting the Large Language Model Inference through Entropy-based N:M Sparsity
by: Li, Yun, et al.
Published: (2023)
by: Li, Yun, et al.
Published: (2023)
Motivating Next-Gen Accelerators with Flexible (N:M) Activation Sparsity via Benchmarking Lightweight Post-Training Sparsification Approaches
by: Alanova, Shirin, et al.
Published: (2025)
by: Alanova, Shirin, et al.
Published: (2025)
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
by: Ramachandran, Akshat, et al.
Published: (2025)
by: Ramachandran, Akshat, et al.
Published: (2025)
Block Sparse Flash Attention
by: Ohayon, Daniel, et al.
Published: (2025)
by: Ohayon, Daniel, et al.
Published: (2025)
TSENOR: Highly-Efficient Algorithm for Finding Transposable N:M Sparse Masks
by: Meng, Xiang, et al.
Published: (2025)
by: Meng, Xiang, et al.
Published: (2025)
PermLLM: Learnable Channel Permutation for N:M Sparse Large Language Models
by: Zou, Lancheng, et al.
Published: (2025)
by: Zou, Lancheng, et al.
Published: (2025)
Stochastic Variance-Reduced Iterative Hard Thresholding in Graph Sparsity Optimization
by: Fox, Derek, et al.
Published: (2024)
by: Fox, Derek, et al.
Published: (2024)
Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers
by: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Published: (2024)
by: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Published: (2024)
Unbiased Gradient Low-Rank Projection
by: Pan, Rui, et al.
Published: (2025)
by: Pan, Rui, et al.
Published: (2025)
Noise-corrected GRPO: From Noisy Rewards to Unbiased Gradients
by: Mansouri, Omar El, et al.
Published: (2025)
by: Mansouri, Omar El, et al.
Published: (2025)
Learning from N-Tuple Data with M Positive Instances: Unbiased Risk Estimation and Theoretical Guarantees
by: Zhang, Miao, et al.
Published: (2025)
by: Zhang, Miao, et al.
Published: (2025)
A Minimum Variance Path Principle for Accurate and Stable Score-Based Density Ratio Estimation
by: Chen, Wei, et al.
Published: (2026)
by: Chen, Wei, et al.
Published: (2026)
On the Rate of Convergence of GD in Non-linear Neural Networks: An Adversarial Robustness Perspective
by: Smorodinsky, Guy, et al.
Published: (2026)
by: Smorodinsky, Guy, et al.
Published: (2026)
On the Optimal Construction of Unbiased Gradient Estimators for Zeroth-Order Optimization
by: Ma, Shaocong, et al.
Published: (2025)
by: Ma, Shaocong, et al.
Published: (2025)
Compositional Sparsity as an Inductive Bias for Neural Architecture Design
by: Lin, Hongyu, et al.
Published: (2026)
by: Lin, Hongyu, et al.
Published: (2026)
Stable Minima Cannot Overfit in Univariate ReLU Networks: Generalization by Large Step Sizes
by: Qiao, Dan, et al.
Published: (2024)
by: Qiao, Dan, et al.
Published: (2024)
Retrieval from Within: An Intrinsic Capability of Attention-Based Models
by: Hoffer, Elad, et al.
Published: (2026)
by: Hoffer, Elad, et al.
Published: (2026)
No Prior, No Leakage: Revisiting Reconstruction Attacks in Trained Neural Networks
by: Refael, Yehonatan, et al.
Published: (2025)
by: Refael, Yehonatan, et al.
Published: (2025)
Class Unbiasing for Generalization in Medical Diagnosis
by: Zuo, Lishi, et al.
Published: (2025)
by: Zuo, Lishi, et al.
Published: (2025)
Time Matters: Scaling Laws for Any Budget
by: Inbar, Itay, et al.
Published: (2024)
by: Inbar, Itay, et al.
Published: (2024)
Revisiting Zeroth-Order Optimization: Minimum-Variance Two-Point Estimators and Directionally Aligned Perturbations
by: Ma, Shaocong, et al.
Published: (2025)
by: Ma, Shaocong, et al.
Published: (2025)
Learning Unbiased Permutations via Flow Matching
by: Min, Yimeng, et al.
Published: (2026)
by: Min, Yimeng, et al.
Published: (2026)
Improving Decision Sparsity
by: Sun, Yiyang, et al.
Published: (2024)
by: Sun, Yiyang, et al.
Published: (2024)
Homeostasis and Sparsity in Transformer
by: Kotyuzanskiy, Leonid, et al.
Published: (2024)
by: Kotyuzanskiy, Leonid, et al.
Published: (2024)
Polar Sparsity: High Throughput Batched LLM Inferencing with Scalable Contextual Sparsity
by: Shrestha, Susav, et al.
Published: (2025)
by: Shrestha, Susav, et al.
Published: (2025)
GNN-VPA: A Variance-Preserving Aggregation Strategy for Graph Neural Networks
by: Schneckenreiter, Lisa, et al.
Published: (2024)
by: Schneckenreiter, Lisa, et al.
Published: (2024)
Weight Concentration Regularization for Improving Pruning Robustness Under High Sparsity
by: Yun, Vincent-Daniel, et al.
Published: (2025)
by: Yun, Vincent-Daniel, et al.
Published: (2025)
Sparsity and Out-of-Distribution Generalization
by: Aaronson, Scott, et al.
Published: (2026)
by: Aaronson, Scott, et al.
Published: (2026)
Sparsity and Superposition in Mixture of Experts
by: Chaudhari, Marmik, et al.
Published: (2025)
by: Chaudhari, Marmik, et al.
Published: (2025)
Similar Items
-
FP4 All the Way: Fully Quantized Training of LLMs
by: Chmiel, Brian, et al.
Published: (2025) -
Scaling FP8 training to trillion-token LLMs
by: Fishman, Maxim, et al.
Published: (2024) -
Normalized Architectures are Natively 4-Bit
by: Fishman, Maxim, et al.
Published: (2026) -
Towards Cheaper Inference in Deep Networks with Lower Bit-Width Accumulators
by: Blumenfeld, Yaniv, et al.
Published: (2024) -
Accurate Neural Training with 4-bit Matrix Multiplications at Standard Formats
by: Chmiel, Brian, et al.
Published: (2021)