Pushing the Limits of Low-Bit Optimizers: A Focus on EMA Dynamics
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Cong, Liang, Wenbin, Yu, Mo, Liu, Anan, Zhang, Ke-Yue, Wang, Shunli, Ma, Lizhuang, Wang, Jianyong, Wang, Jun, Zhang, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Graph-enhanced Optimizers for Structure-aware Recommendation Embedding Evolution
by: Xu, Cong, et al.
Published: (2023)
by: Xu, Cong, et al.
Published: (2023)
CLAQ: Pushing the Limits of Low-Bit Post-Training Quantization for LLMs
by: Wang, Haoyu, et al.
Published: (2024)
by: Wang, Haoyu, et al.
Published: (2024)
Are LLM-based Recommenders Already the Best? Simple Scaled Cross-entropy Unleashes the Potential of Traditional Sequential Recommenders
by: Xu, Cong, et al.
Published: (2024)
by: Xu, Cong, et al.
Published: (2024)
PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training Quantization Methods for Large Language Models
by: Zhao, Jiaqi, et al.
Published: (2025)
by: Zhao, Jiaqi, et al.
Published: (2025)
Understanding the Role of Cross-Entropy Loss in Fairly Evaluating Large Language Model-based Recommendation
by: Xu, Cong, et al.
Published: (2024)
by: Xu, Cong, et al.
Published: (2024)
1-Bit FQT: Pushing the Limit of Fully Quantized Training to 1-bit
by: Gao, Chang, et al.
Published: (2024)
by: Gao, Chang, et al.
Published: (2024)
BitHEP -- The Limits of Low-Precision ML in HEP
by: Krause, Claudius, et al.
Published: (2025)
by: Krause, Claudius, et al.
Published: (2025)
PlotCraft: Pushing the Limits of LLMs for Complex and Interactive Data Visualization
by: Zhang, Jiajun, et al.
Published: (2025)
by: Zhang, Jiajun, et al.
Published: (2025)
Pushing The Limit of LLM Capacity for Text Classification
by: Zhang, Yazhou, et al.
Published: (2024)
by: Zhang, Yazhou, et al.
Published: (2024)
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache
by: Sharma, Akshat, et al.
Published: (2024)
by: Sharma, Akshat, et al.
Published: (2024)
FocusLLM: Precise Understanding of Long Context by Dynamic Condensing
by: Li, Zhenyu, et al.
Published: (2024)
by: Li, Zhenyu, et al.
Published: (2024)
BiDM: Pushing the Limit of Quantization for Diffusion Models
by: Zheng, Xingyu, et al.
Published: (2024)
by: Zheng, Xingyu, et al.
Published: (2024)
CRTrack: Low-Light Semi-Supervised Multi-object Tracking Based on Consistency Regularization
by: Zhao, Zijing, et al.
Published: (2025)
by: Zhao, Zijing, et al.
Published: (2025)
EMA: Efficient Model Adaptation for Learning-based Systems
by: Yu, Daiyang, et al.
Published: (2026)
by: Yu, Daiyang, et al.
Published: (2026)
In situ Implanting 3D Carbon Network Reinforced Zinc Composite by Powder Metallurgy for Highly Reversible Zn‐based Battery Anodes
by: Jingxian Wang, et al.
Published: (2024)
by: Jingxian Wang, et al.
Published: (2024)
In situ Implanting 3D Carbon Network Reinforced Zinc Composite by Powder Metallurgy for Highly Reversible Zn‐based Battery Anodes
by: Jingxian Wang, et al.
Published: (2024)
by: Jingxian Wang, et al.
Published: (2024)
Pangu Ultra: Pushing the Limits of Dense Large Language Models on Ascend NPUs
by: Yin, Yichun, et al.
Published: (2025)
by: Yin, Yichun, et al.
Published: (2025)
Pore‐Matched Sponge for Microorganisms Pushes Electron Extraction Limit in Microbial Fuel Cells (Small 7/2024)
by: Ke Feng, et al.
Published: (2024)
by: Ke Feng, et al.
Published: (2024)
An adaptive weighted memory event‐triggered scheme for switched positive systems
by: Shunli Zhao, et al.
Published: (2025)
by: Shunli Zhao, et al.
Published: (2025)
GaitAdapt: Continual Learning for Evolving Gait Recognition
by: Wang, Jingjie, et al.
Published: (2025)
by: Wang, Jingjie, et al.
Published: (2025)
EMA Policy Gradient: Taming Reinforcement Learning for LLMs with EMA Anchor and Top-k KL
by: Zhang, Lunjun, et al.
Published: (2026)
by: Zhang, Lunjun, et al.
Published: (2026)
Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025
by: Ying, Zonghao, et al.
Published: (2025)
by: Ying, Zonghao, et al.
Published: (2025)
BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
by: Xin, Detai, et al.
Published: (2024)
by: Xin, Detai, et al.
Published: (2024)
T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on Edge
by: Wei, Jianyu, et al.
Published: (2024)
by: Wei, Jianyu, et al.
Published: (2024)
SaRA: High-Efficient Diffusion Model Fine-tuning with Progressive Sparse Low-Rank Adaptation
by: Hu, Teng, et al.
Published: (2024)
by: Hu, Teng, et al.
Published: (2024)
InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
by: Li, Ke, et al.
Published: (2026)
by: Li, Ke, et al.
Published: (2026)
QVGen: Pushing the Limit of Quantized Video Generative Models
by: Huang, Yushi, et al.
Published: (2025)
by: Huang, Yushi, et al.
Published: (2025)
Learning Interpretable Rules for Scalable Data Representation and Classification
by: Wang, Zhuo, et al.
Published: (2023)
by: Wang, Zhuo, et al.
Published: (2023)
Revealing Vulnerabilities in Stable Diffusion via Targeted Attacks
by: Zhang, Chenyu, et al.
Published: (2024)
by: Zhang, Chenyu, et al.
Published: (2024)
SSR: Pushing the Limit of Spatial Intelligence with Structured Scene Reasoning
by: Zhang, Yi, et al.
Published: (2026)
by: Zhang, Yi, et al.
Published: (2026)
MaskFocus: Focusing Policy Optimization on Critical Steps for Masked Image Generation
by: Zhang, Guohui, et al.
Published: (2025)
by: Zhang, Guohui, et al.
Published: (2025)
AdR-Gaussian: Accelerating Gaussian Splatting with Adaptive Radius
by: Wang, Xinzhe, et al.
Published: (2024)
by: Wang, Xinzhe, et al.
Published: (2024)
Real-IAD Variety: Pushing Industrial Anomaly Detection Dataset to a Modern Era
by: Zhu, Wenbing, et al.
Published: (2025)
by: Zhu, Wenbing, et al.
Published: (2025)
CAdam: Confidence-Based Optimization for Online Learning
by: Wang, Shaowen, et al.
Published: (2024)
by: Wang, Shaowen, et al.
Published: (2024)
Revealing the Two-Fold Ambiguity: Tau Momentum Reconstruction and Its Impact on Entanglement Observables
by: Zhou, Xiang, et al.
Published: (2026)
by: Zhou, Xiang, et al.
Published: (2026)
Shedding the Bits: Pushing the Boundaries of Quantization with Minifloats on FPGAs
by: Aggarwal, Shivam, et al.
Published: (2023)
by: Aggarwal, Shivam, et al.
Published: (2023)
DAGPrompT: Pushing the Limits of Graph Prompting with a Distribution-aware Graph Prompt Tuning Approach
by: Chen, Qin, et al.
Published: (2025)
by: Chen, Qin, et al.
Published: (2025)
Class-Imbalanced Semi-Supervised Learning for Large-Scale Point Cloud Semantic Segmentation via Decoupling Optimization
by: Li, Mengtian, et al.
Published: (2024)
by: Li, Mengtian, et al.
Published: (2024)
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
by: Shao, Zhihong, et al.
Published: (2024)
by: Shao, Zhihong, et al.
Published: (2024)
Wideband and High‐Gain Folded Transmit‐Array Antenna Based on 3‐Bit FPTP Metasurface
by: Qiang Wang, et al.
Published: (2025)
by: Qiang Wang, et al.
Published: (2025)
Similar Items
-
Graph-enhanced Optimizers for Structure-aware Recommendation Embedding Evolution
by: Xu, Cong, et al.
Published: (2023) -
CLAQ: Pushing the Limits of Low-Bit Post-Training Quantization for LLMs
by: Wang, Haoyu, et al.
Published: (2024) -
Are LLM-based Recommenders Already the Best? Simple Scaled Cross-entropy Unleashes the Potential of Traditional Sequential Recommenders
by: Xu, Cong, et al.
Published: (2024) -
PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training Quantization Methods for Large Language Models
by: Zhao, Jiaqi, et al.
Published: (2025) -
Understanding the Role of Cross-Entropy Loss in Fairly Evaluating Large Language Model-based Recommendation
by: Xu, Cong, et al.
Published: (2024)