4-bit Shampoo for Memory-Efficient Network Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Sike, Zhou, Pan, Li, Jia, Huang, Hua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Zeroth-Order Fine-Tuning of LLMs in Random Subspaces
von: Yu, Ziming, et al.
Veröffentlicht: (2024)
von: Yu, Ziming, et al.
Veröffentlicht: (2024)
Purifying Shampoo: Investigating Shampoo's Heuristics by Decomposing its Preconditioner
von: Eschenhagen, Runa, et al.
Veröffentlicht: (2025)
von: Eschenhagen, Runa, et al.
Veröffentlicht: (2025)
SOAP: Improving and Stabilizing Shampoo using Adam
von: Vyas, Nikhil, et al.
Veröffentlicht: (2024)
von: Vyas, Nikhil, et al.
Veröffentlicht: (2024)
Clarifying Shampoo: Adapting Spectral Descent to Stochasticity and the Parameter Trajectory
von: Eschenhagen, Runa, et al.
Veröffentlicht: (2026)
von: Eschenhagen, Runa, et al.
Veröffentlicht: (2026)
Pro-KLShampoo: Projected KL-Shampoo with Whitening Recovered by Orthogonalization
von: Sun, Ruotong, et al.
Veröffentlicht: (2026)
von: Sun, Ruotong, et al.
Veröffentlicht: (2026)
FOAM: Blocked State Folding for Memory-Efficient LLM Training
von: Wen, Ziqing, et al.
Veröffentlicht: (2025)
von: Wen, Ziqing, et al.
Veröffentlicht: (2025)
SpanGNN: Towards Memory-Efficient Graph Neural Networks via Spanning Subgraph Training
von: Gu, Xizhi, et al.
Veröffentlicht: (2024)
von: Gu, Xizhi, et al.
Veröffentlicht: (2024)
NeuZip: Memory-Efficient Training and Inference with Dynamic Compression of Neural Networks
von: Hao, Yongchang, et al.
Veröffentlicht: (2024)
von: Hao, Yongchang, et al.
Veröffentlicht: (2024)
FlashOptim: Optimizers for Memory-Efficient Training
von: Ortiz, Jose Javier Gonzalez, et al.
Veröffentlicht: (2026)
von: Ortiz, Jose Javier Gonzalez, et al.
Veröffentlicht: (2026)
Kitty: Accurate and Efficient 2-bit KV Cache Quantization with Dynamic Channel-wise Precision Boost
von: Xia, Haojun, et al.
Veröffentlicht: (2025)
von: Xia, Haojun, et al.
Veröffentlicht: (2025)
4bit-Quantization in Vector-Embedding for RAG
von: Jeong, Taehee
Veröffentlicht: (2025)
von: Jeong, Taehee
Veröffentlicht: (2025)
StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs
von: Luo, Qijun, et al.
Veröffentlicht: (2025)
von: Luo, Qijun, et al.
Veröffentlicht: (2025)
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
von: Yan, Xianglong, et al.
Veröffentlicht: (2026)
von: Yan, Xianglong, et al.
Veröffentlicht: (2026)
FOAM: Frequency and Operator Error-Based Adaptive Damping Method for Reducing Staleness-Oriented Error for Shampoo
von: Nam, Kyunghun, et al.
Veröffentlicht: (2026)
von: Nam, Kyunghun, et al.
Veröffentlicht: (2026)
Low-bit Model Quantization for Deep Neural Networks: A Survey
von: Liu, Kai, et al.
Veröffentlicht: (2025)
von: Liu, Kai, et al.
Veröffentlicht: (2025)
any4: Learned 4-bit Numeric Representation for LLMs
von: Elhoushi, Mostafa, et al.
Veröffentlicht: (2025)
von: Elhoushi, Mostafa, et al.
Veröffentlicht: (2025)
BEND: Bagging Deep Learning Training Based on Efficient Neural Network Diffusion
von: Wei, Jia, et al.
Veröffentlicht: (2024)
von: Wei, Jia, et al.
Veröffentlicht: (2024)
A Survey on Memory-Efficient Transformer-Based Model Training in AI for Science
von: Tian, Kaiyuan, et al.
Veröffentlicht: (2025)
von: Tian, Kaiyuan, et al.
Veröffentlicht: (2025)
Train Small, Infer Large: Memory-Efficient LoRA Training for Large Language Models
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
Continual Quantization-Aware Pre-Training: When to transition from 16-bit to 1.58-bit pre-training for BitNet language models?
von: Nielsen, Jacob, et al.
Veröffentlicht: (2025)
von: Nielsen, Jacob, et al.
Veröffentlicht: (2025)
Conda: Column-Normalized Adam for Training Large Language Models Faster
von: Wang, Junjie, et al.
Veröffentlicht: (2025)
von: Wang, Junjie, et al.
Veröffentlicht: (2025)
FedProphet: Memory-Efficient Federated Adversarial Training via Robust and Consistent Cascade Learning
von: Tang, Minxue, et al.
Veröffentlicht: (2024)
von: Tang, Minxue, et al.
Veröffentlicht: (2024)
SlimPipe: Memory-Thrifty and Efficient Pipeline Parallelism for Long-Context LLM Training
von: Li, Zhouyang, et al.
Veröffentlicht: (2025)
von: Li, Zhouyang, et al.
Veröffentlicht: (2025)
MCUBERT: Memory-Efficient BERT Inference on Commodity Microcontrollers
von: Yang, Zebin, et al.
Veröffentlicht: (2024)
von: Yang, Zebin, et al.
Veröffentlicht: (2024)
Gradient Weight-normalized Low-rank Projection for Efficient LLM Training
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024)
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024)
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling
von: Ji, Xiaodong, et al.
Veröffentlicht: (2025)
von: Ji, Xiaodong, et al.
Veröffentlicht: (2025)
Data Warmup: Complexity-Aware Curricula for Efficient Diffusion Training
von: Lin, Jinhong, et al.
Veröffentlicht: (2026)
von: Lin, Jinhong, et al.
Veröffentlicht: (2026)
VQ4ALL: Efficient Neural Network Representation via a Universal Codebook
von: Deng, Juncan, et al.
Veröffentlicht: (2024)
von: Deng, Juncan, et al.
Veröffentlicht: (2024)
COAT: Compressing Optimizer states and Activation for Memory-Efficient FP8 Training
von: Xi, Haocheng, et al.
Veröffentlicht: (2024)
von: Xi, Haocheng, et al.
Veröffentlicht: (2024)
Memory-Efficient LLM Training with Dynamic Sparsity: From Stability to Practical Scaling
von: Xiao, Qiao, et al.
Veröffentlicht: (2026)
von: Xiao, Qiao, et al.
Veröffentlicht: (2026)
ssProp: Energy-Efficient Training for Convolutional Neural Networks with Scheduled Sparse Back Propagation
von: Zhong, Lujia, et al.
Veröffentlicht: (2024)
von: Zhong, Lujia, et al.
Veröffentlicht: (2024)
UltraEdit: Training-, Subject-, and Memory-Free Lifelong Editing in Language Models
von: Gu, Xiaojie, et al.
Veröffentlicht: (2025)
von: Gu, Xiaojie, et al.
Veröffentlicht: (2025)
CoMERA: Computing- and Memory-Efficient Training via Rank-Adaptive Tensor Optimization
von: Yang, Zi, et al.
Veröffentlicht: (2024)
von: Yang, Zi, et al.
Veröffentlicht: (2024)
ICQuant: Index Coding enables Low-bit LLM Quantization
von: Li, Xinlin, et al.
Veröffentlicht: (2025)
von: Li, Xinlin, et al.
Veröffentlicht: (2025)
Entropy-Guided Data-Efficient Training for Multimodal Reasoning Reward Models
von: Yang, Shidong, et al.
Veröffentlicht: (2026)
von: Yang, Shidong, et al.
Veröffentlicht: (2026)
SageBwd: A Trainable Low-bit Attention
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
The Hidden Power of Pure 16-bit Floating-Point Neural Networks
von: Yun, Juyoung, et al.
Veröffentlicht: (2023)
von: Yun, Juyoung, et al.
Veröffentlicht: (2023)
IDInit: A Universal and Stable Initialization Method for Neural Network Training
von: Pan, Yu, et al.
Veröffentlicht: (2025)
von: Pan, Yu, et al.
Veröffentlicht: (2025)
GSR-GNN: Training Acceleration and Memory-Saving Framework of Deep GNNs on Circuit Graph
von: Luo, Yuebo, et al.
Veröffentlicht: (2026)
von: Luo, Yuebo, et al.
Veröffentlicht: (2026)
ProTrain: Efficient LLM Training via Memory-Aware Techniques
von: Yang, Hanmei, et al.
Veröffentlicht: (2024)
von: Yang, Hanmei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Zeroth-Order Fine-Tuning of LLMs in Random Subspaces
von: Yu, Ziming, et al.
Veröffentlicht: (2024) -
Purifying Shampoo: Investigating Shampoo's Heuristics by Decomposing its Preconditioner
von: Eschenhagen, Runa, et al.
Veröffentlicht: (2025) -
SOAP: Improving and Stabilizing Shampoo using Adam
von: Vyas, Nikhil, et al.
Veröffentlicht: (2024) -
Clarifying Shampoo: Adapting Spectral Descent to Stochasticity and the Parameter Trajectory
von: Eschenhagen, Runa, et al.
Veröffentlicht: (2026) -
Pro-KLShampoo: Projected KL-Shampoo with Whitening Recovered by Orthogonalization
von: Sun, Ruotong, et al.
Veröffentlicht: (2026)