Understanding Quantization of Optimizer States in LLM Pre-training: Dynamics of State Staleness and Effectiveness of State Resets
Fuente:
arXiv
Saved in:
| Main Authors: | Topollai, Kristi, Choromanska, Anna |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adaptive Memory Momentum via a Model-Based Framework for Deep Learning Optimization
by: Topollai, Kristi, et al.
Published: (2025)
by: Topollai, Kristi, et al.
Published: (2025)
Task-Level Contrastiveness for Cross-Domain Few-Shot Learning
by: Topollai, Kristi, et al.
Published: (2025)
by: Topollai, Kristi, et al.
Published: (2025)
Worker Disagreement Reveals Sharp Directions in Local SGD
by: Dimlioglu, Tolga, et al.
Published: (2026)
by: Dimlioglu, Tolga, et al.
Published: (2026)
Outer-Momentum Restarting in High-Dimensional Two-Phase Optimization
by: Topollai, Kristi, et al.
Published: (2026)
by: Topollai, Kristi, et al.
Published: (2026)
OncoReason: Structuring Clinical Reasoning in LLMs for Robust and Interpretable Survival Prediction
by: Hemadri, Raghu Vamshi, et al.
Published: (2025)
by: Hemadri, Raghu Vamshi, et al.
Published: (2025)
Effective Quantization of Muon Optimizer States
by: Gupta, Aman, et al.
Published: (2025)
by: Gupta, Aman, et al.
Published: (2025)
Streamlining Industrial Contract Management with Retrieval-Augmented LLMs
by: Topollai, Kristi, et al.
Published: (2025)
by: Topollai, Kristi, et al.
Published: (2025)
Self-Supervised Representation Learning with Joint Embedding Predictive Architecture for Automotive LiDAR Object Detection
by: Zhu, Haoran, et al.
Published: (2025)
by: Zhu, Haoran, et al.
Published: (2025)
A Survey of Optimization Methods for Training DL Models: Theoretical Perspective on Convergence and Generalization
by: Wang, Jing, et al.
Published: (2025)
by: Wang, Jing, et al.
Published: (2025)
Communication-Efficient Distributed Training for Collaborative Flat Optima Recovery in Deep Learning
by: Dimlioglu, Tolga, et al.
Published: (2025)
by: Dimlioglu, Tolga, et al.
Published: (2025)
This Too Shall Pass: Removing Stale Observations in Dynamic Bayesian Optimization
by: Bardou, Anthony, et al.
Published: (2024)
by: Bardou, Anthony, et al.
Published: (2024)
Adjacent Leader Decentralized Stochastic Gradient Descent
by: He, Haoze, et al.
Published: (2024)
by: He, Haoze, et al.
Published: (2024)
The Fourth State: Signed-Zero Ternary for Stable LLM Quantization (and More)
by: Uhlmann, Jeffrey
Published: (2025)
by: Uhlmann, Jeffrey
Published: (2025)
GRAWA: Gradient-based Weighted Averaging for Distributed Training of Deep Learning Models
by: Dimlioglu, Tolga, et al.
Published: (2024)
by: Dimlioglu, Tolga, et al.
Published: (2024)
Priming: Hybrid State Space Models From Pre-trained Transformers
by: Chattopadhyay, Aditya, et al.
Published: (2026)
by: Chattopadhyay, Aditya, et al.
Published: (2026)
Sample-efficient LLM Optimization with Reset Replay
by: Liu, Zichuan, et al.
Published: (2025)
by: Liu, Zichuan, et al.
Published: (2025)
TAME: Task Agnostic Continual Learning using Multiple Experts
by: Zhu, Haoran, et al.
Published: (2022)
by: Zhu, Haoran, et al.
Published: (2022)
Accelerating Recommender Model Training by Dynamically Skipping Stale Embeddings
by: Maboud, Yassaman Ebrahimzadeh, et al.
Published: (2024)
by: Maboud, Yassaman Ebrahimzadeh, et al.
Published: (2024)
Quantifying Memory Utilization with Effective State-Size
by: Parnichkun, Rom N., et al.
Published: (2025)
by: Parnichkun, Rom N., et al.
Published: (2025)
State Regularized Policy Optimization on Data with Dynamics Shift
by: Xue, Zhenghai, et al.
Published: (2023)
by: Xue, Zhenghai, et al.
Published: (2023)
MUON+: Towards More Effective Muon via One Additional Normalization Step for LLM Pre-training
by: Zhang, Ruijie, et al.
Published: (2026)
by: Zhang, Ruijie, et al.
Published: (2026)
Discrete Semantic States and Hamiltonian Dynamics in LLM Embedding Spaces
by: Laine, Timo Aukusti
Published: (2025)
by: Laine, Timo Aukusti
Published: (2025)
Quamba2: A Robust and Scalable Post-training Quantization Framework for Selective State Space Models
by: Chiang, Hung-Yueh, et al.
Published: (2025)
by: Chiang, Hung-Yueh, et al.
Published: (2025)
StateX: Enhancing RNN Recall via Post-training State Expansion
by: Shen, Xingyu, et al.
Published: (2025)
by: Shen, Xingyu, et al.
Published: (2025)
Zero-Shot Cross-City Generalization in End-to-End Autonomous Driving: Self-Supervised versus Supervised Representations
by: Naeinian, Fatemeh, et al.
Published: (2026)
by: Naeinian, Fatemeh, et al.
Published: (2026)
Quantizing Small-Scale State-Space Models for Edge AI
by: Zhao, Leo, et al.
Published: (2025)
by: Zhao, Leo, et al.
Published: (2025)
Pre-trained Gaussian Processes for Bayesian Optimization
by: Wang, Zi, et al.
Published: (2021)
by: Wang, Zi, et al.
Published: (2021)
Latent State Models of Training Dynamics
by: Hu, Michael Y., et al.
Published: (2023)
by: Hu, Michael Y., et al.
Published: (2023)
FedStaleWeight: Buffered Asynchronous Federated Learning with Fair Aggregation via Staleness Reweighting
by: Ma, Jeffrey, et al.
Published: (2024)
by: Ma, Jeffrey, et al.
Published: (2024)
Absolute State-wise Constrained Policy Optimization: High-Probability State-wise Constraints Satisfaction
by: Zhao, Weiye, et al.
Published: (2024)
by: Zhao, Weiye, et al.
Published: (2024)
Cosine-Gated Adam-Decay: Drop-In Staleness-Aware Outer Optimization for Decoupled DiLoCo
by: Shah, Vatsal, et al.
Published: (2026)
by: Shah, Vatsal, et al.
Published: (2026)
State-wise Constrained Policy Optimization
by: Zhao, Weiye, et al.
Published: (2023)
by: Zhao, Weiye, et al.
Published: (2023)
Tahakom LLM Guidelines and Recipes: From Pre-training Data to an Arabic LLM
by: AlOtaibi, Areej, et al.
Published: (2025)
by: AlOtaibi, Areej, et al.
Published: (2025)
Communication Efficient LLM Pre-training with SparseLoCo
by: Sarfi, Amir, et al.
Published: (2025)
by: Sarfi, Amir, et al.
Published: (2025)
Q-S5: Towards Quantized State Space Models
by: Abreu, Steven, et al.
Published: (2024)
by: Abreu, Steven, et al.
Published: (2024)
GeoDynamics: A Geometric State-Space Neural Network for Understanding Brain Dynamics on Riemannian Manifolds
by: Dan, Tingting, et al.
Published: (2026)
by: Dan, Tingting, et al.
Published: (2026)
Markovian Circuit Tracing for Transformer State Dynamic
by: X, Abdullah
Published: (2026)
by: X, Abdullah
Published: (2026)
Dynamical Survival Analysis with Controlled Latent States
by: Bleistein, Linus, et al.
Published: (2024)
by: Bleistein, Linus, et al.
Published: (2024)
Dataset Reset Policy Optimization for RLHF
by: Chang, Jonathan D., et al.
Published: (2024)
by: Chang, Jonathan D., et al.
Published: (2024)
Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning
by: Chen, Zizhe, et al.
Published: (2026)
by: Chen, Zizhe, et al.
Published: (2026)
Similar Items
-
Adaptive Memory Momentum via a Model-Based Framework for Deep Learning Optimization
by: Topollai, Kristi, et al.
Published: (2025) -
Task-Level Contrastiveness for Cross-Domain Few-Shot Learning
by: Topollai, Kristi, et al.
Published: (2025) -
Worker Disagreement Reveals Sharp Directions in Local SGD
by: Dimlioglu, Tolga, et al.
Published: (2026) -
Outer-Momentum Restarting in High-Dimensional Two-Phase Optimization
by: Topollai, Kristi, et al.
Published: (2026) -
OncoReason: Structuring Clinical Reasoning in LLMs for Robust and Interpretable Survival Prediction
by: Hemadri, Raghu Vamshi, et al.
Published: (2025)