Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Cheng, Wenhua, Zhang, Weiwei, Shen, Haihao, Cai, Yiyang, He, Xin, Lv, Kaokao, Liu, Yi |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SignRoundV2: Toward Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMs
by: Cheng, Wenhua, et al.
Published: (2025)
by: Cheng, Wenhua, et al.
Published: (2025)
Efficient Post-training Quantization with FP8 Formats
by: Shen, Haihao, et al.
Published: (2023)
by: Shen, Haihao, et al.
Published: (2023)
Round Attention: A Novel Round-Level Attention Mechanism to Accelerate LLM Inference
by: Tang, Yaohua, et al.
Published: (2025)
by: Tang, Yaohua, et al.
Published: (2025)
GWQ: Gradient-Aware Weight Quantization for Large Language Models
by: Shao, Yihua, et al.
Published: (2024)
by: Shao, Yihua, et al.
Published: (2024)
RoLoRA: Fine-tuning Rotated Outlier-free LLMs for Effective Weight-Activation Quantization
by: Huang, Xijie, et al.
Published: (2024)
by: Huang, Xijie, et al.
Published: (2024)
Do pretrained Transformers Learn In-Context by Gradient Descent?
by: Shen, Lingfeng, et al.
Published: (2023)
by: Shen, Lingfeng, et al.
Published: (2023)
PlatoLM: Teaching LLMs in Multi-Round Dialogue via a User Simulator
by: Kong, Chuyi, et al.
Published: (2023)
by: Kong, Chuyi, et al.
Published: (2023)
CDQuant: Greedy Coordinate Descent for Accurate LLM Quantization
by: Nair, Pranav Ajit, et al.
Published: (2024)
by: Nair, Pranav Ajit, et al.
Published: (2024)
Quantization Meets dLLMs: A Systematic Study of Post-training Quantization for Diffusion LLMs
by: Lin, Haokun, et al.
Published: (2025)
by: Lin, Haokun, et al.
Published: (2025)
Enhanced Sign Language Translation between American Sign Language (ASL) and Indian Sign Language (ISL) Using LLMs
by: Kumar, Malay, et al.
Published: (2024)
by: Kumar, Malay, et al.
Published: (2024)
AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs
by: He, Di, et al.
Published: (2025)
by: He, Di, et al.
Published: (2025)
ProtRLSearch: A Multi-Round Multimodal Protein Search Agent with Large Language Models Trained via Reinforcement Learning
by: Liu, Congying, et al.
Published: (2026)
by: Liu, Congying, et al.
Published: (2026)
POPI: Personalizing LLMs via Optimized Natural Language Preference Inference
by: Chen, Yizhuo, et al.
Published: (2025)
by: Chen, Yizhuo, et al.
Published: (2025)
Catastrophic Failure of LLM Unlearning via Quantization
by: Zhang, Zhiwei, et al.
Published: (2024)
by: Zhang, Zhiwei, et al.
Published: (2024)
Rapid Optimization for Jailbreaking LLMs via Subconscious Exploitation and Echopraxia
by: Shen, Guangyu, et al.
Published: (2024)
by: Shen, Guangyu, et al.
Published: (2024)
Router-R1: Teaching LLMs Multi-Round Routing and Aggregation via Reinforcement Learning
by: Zhang, Haozhen, et al.
Published: (2025)
by: Zhang, Haozhen, et al.
Published: (2025)
ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs
by: Chen, Justin Chih-Yao, et al.
Published: (2023)
by: Chen, Justin Chih-Yao, et al.
Published: (2023)
Fitting Is Not Enough: Smoothness in Extremely Quantized LLMs
by: Xu, Yuzhuang, et al.
Published: (2026)
by: Xu, Yuzhuang, et al.
Published: (2026)
Continuous Approximations for Improving Quantization Aware Training of LLMs
by: Li, He, et al.
Published: (2024)
by: Li, He, et al.
Published: (2024)
SoLoPO: Unlocking Long-Context Capabilities in LLMs via Short-to-Long Preference Optimization
by: Sun, Huashan, et al.
Published: (2025)
by: Sun, Huashan, et al.
Published: (2025)
SSL-SSAW: Self-Supervised Learning with Sigmoid Self-Attention Weighting for Question-Based Sign Language Translation
by: Liu, Zekang, et al.
Published: (2025)
by: Liu, Zekang, et al.
Published: (2025)
Towards Self-Robust LLMs: Intrinsic Prompt Noise Resistance via CoIPO
by: Yang, Xin, et al.
Published: (2026)
by: Yang, Xin, et al.
Published: (2026)
IOPO: Empowering LLMs with Complex Instruction Following via Input-Output Preference Optimization
by: Zhang, Xinghua, et al.
Published: (2024)
by: Zhang, Xinghua, et al.
Published: (2024)
GradPruner: Gradient-Guided Layer Pruning Enabling Efficient Fine-Tuning and Inference for LLMs
by: Huang, Wei, et al.
Published: (2026)
by: Huang, Wei, et al.
Published: (2026)
Klear-Reasoner: Advancing Reasoning Capability via Gradient-Preserving Clipping Policy Optimization
by: Su, Zhenpeng, et al.
Published: (2025)
by: Su, Zhenpeng, et al.
Published: (2025)
Interpreting the Effects of Quantization on LLMs
by: Singh, Manpreet, et al.
Published: (2025)
by: Singh, Manpreet, et al.
Published: (2025)
Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space
by: Zhang, Zhen, et al.
Published: (2025)
by: Zhang, Zhen, et al.
Published: (2025)
Causal Intersectionality and Dual Form of Gradient Descent for Multimodal Analysis: a Case Study on Hateful Memes
by: Miyanishi, Yosuke, et al.
Published: (2023)
by: Miyanishi, Yosuke, et al.
Published: (2023)
Online Causal Kalman Filtering for Stable and Effective Policy Optimization
by: He, Shuo, et al.
Published: (2026)
by: He, Shuo, et al.
Published: (2026)
GreenTEA: Gradient Descent with Topic-modeling and Evolutionary Auto-prompting
by: Dong, Zheng, et al.
Published: (2025)
by: Dong, Zheng, et al.
Published: (2025)
CompactQE: Interpretable Translation Quality Estimation via Small Open-Weight LLMs
by: Guttmann, Kamil, et al.
Published: (2026)
by: Guttmann, Kamil, et al.
Published: (2026)
RSQ: Learning from Important Tokens Leads to Better Quantized LLMs
by: Sung, Yi-Lin, et al.
Published: (2025)
by: Sung, Yi-Lin, et al.
Published: (2025)
BATQuant: Outlier-resilient MXFP4 Quantization via Learnable Block-wise Optimization
by: Li, Ji-Fu, et al.
Published: (2026)
by: Li, Ji-Fu, et al.
Published: (2026)
Compensate Quantization Errors: Make Weights Hierarchical to Compensate Each Other
by: Gao, Yifei, et al.
Published: (2024)
by: Gao, Yifei, et al.
Published: (2024)
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
by: Su, Zunhai, et al.
Published: (2025)
by: Su, Zunhai, et al.
Published: (2025)
RankLLM: Weighted Ranking of LLMs by Quantifying Question Difficulty
by: Zhang, Ziqian, et al.
Published: (2026)
by: Zhang, Ziqian, et al.
Published: (2026)
Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge
by: Shen, Xuan, et al.
Published: (2023)
by: Shen, Xuan, et al.
Published: (2023)
Improving Low-Resource Machine Translation via Round-Trip Reinforcement Learning
by: Attia, Ahmed, et al.
Published: (2026)
by: Attia, Ahmed, et al.
Published: (2026)
Middo: Model-Informed Dynamic Data Optimization for Enhanced LLM Fine-Tuning via Closed-Loop Learning
by: Tang, Zinan, et al.
Published: (2025)
by: Tang, Zinan, et al.
Published: (2025)
When Personalization Misleads: Understanding and Mitigating Hallucinations in Personalized LLMs
by: Sun, Zhongxiang, et al.
Published: (2026)
by: Sun, Zhongxiang, et al.
Published: (2026)
Similar Items
-
SignRoundV2: Toward Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMs
by: Cheng, Wenhua, et al.
Published: (2025) -
Efficient Post-training Quantization with FP8 Formats
by: Shen, Haihao, et al.
Published: (2023) -
Round Attention: A Novel Round-Level Attention Mechanism to Accelerate LLM Inference
by: Tang, Yaohua, et al.
Published: (2025) -
GWQ: Gradient-Aware Weight Quantization for Large Language Models
by: Shao, Yihua, et al.
Published: (2024) -
RoLoRA: Fine-tuning Rotated Outlier-free LLMs for Effective Weight-Activation Quantization
by: Huang, Xijie, et al.
Published: (2024)