One QuantLLM for ALL: Fine-tuning Quantized LLMs Once for Efficient Deployments
Fuente:
arXiv
Saved in:
| Main Authors: | Yi, Ke, Xu, Yuhui, Chang, Heng, Tang, Chen, Meng, Yuan, Zhang, Tong, Li, Jia |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EasyQuant: An Efficient Data-free Quantization Algorithm for LLMs
by: Tang, Hanlin, et al.
Published: (2024)
by: Tang, Hanlin, et al.
Published: (2024)
Detect an Object At Once without Fine-tuning
by: Hao, Junyu, et al.
Published: (2024)
by: Hao, Junyu, et al.
Published: (2024)
InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
by: Li, Ke, et al.
Published: (2026)
by: Li, Ke, et al.
Published: (2026)
SliderQuant: Accurate Post-Training Quantization for LLMs
by: Wang, Shigeng, et al.
Published: (2026)
by: Wang, Shigeng, et al.
Published: (2026)
Quantized Evolution Strategies: High-precision Fine-tuning of Quantized LLMs at Low-precision Cost
by: Xu, Yinggan, et al.
Published: (2026)
by: Xu, Yinggan, et al.
Published: (2026)
PrivQuant: Communication-Efficient Private Inference with Quantized Network/Protocol Co-Optimization
by: Xu, Tianshi, et al.
Published: (2024)
by: Xu, Tianshi, et al.
Published: (2024)
Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
by: Zhang, Tuo, et al.
Published: (2025)
by: Zhang, Tuo, et al.
Published: (2025)
Quant.npu: Enabling Efficient Mobile NPU Inference for on-device LLMs via Fully Static Quantization
by: Zhang, Jinghe, et al.
Published: (2026)
by: Zhang, Jinghe, et al.
Published: (2026)
DuQuant++: Fine-grained Rotation Enhances Microscaling FP4 Quantization
by: Lin, Haokun, et al.
Published: (2026)
by: Lin, Haokun, et al.
Published: (2026)
DFQ-ViT: Data-Free Quantization for Vision Transformers without Fine-tuning
by: Tong, Yujia, et al.
Published: (2025)
by: Tong, Yujia, et al.
Published: (2025)
NestQuant: Nested Lattice Quantization for Matrix Products and LLMs
by: Savkin, Semyon, et al.
Published: (2025)
by: Savkin, Semyon, et al.
Published: (2025)
Q-realign: Piggybacking Realignment on Quantization for Safe and Efficient LLM Deployment
by: Tan, Qitao, et al.
Published: (2026)
by: Tan, Qitao, et al.
Published: (2026)
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
by: Xu, Zukang, et al.
Published: (2025)
by: Xu, Zukang, et al.
Published: (2025)
On-the-Fly Adaptation to Quantization: Configuration-Aware LoRA for Efficient Fine-Tuning of Quantized LLMs
by: Ye, Rongguang, et al.
Published: (2025)
by: Ye, Rongguang, et al.
Published: (2025)
QuantLRM: Quantization of Large Reasoning Models via Fine-Tuning Signals
by: Zhang, Nan, et al.
Published: (2026)
by: Zhang, Nan, et al.
Published: (2026)
ALL-FEM: Agentic Large Language models Fine-tuned for Finite Element Methods
by: Deotale, Rushikesh, et al.
Published: (2026)
by: Deotale, Rushikesh, et al.
Published: (2026)
ZorBA: Zeroth-order Federated Fine-tuning of LLMs with Heterogeneous Block Activation
by: Meng, Chuiyang, et al.
Published: (2026)
by: Meng, Chuiyang, et al.
Published: (2026)
DilateQuant: Accurate and Efficient Diffusion Quantization via Weight Dilation
by: Liu, Xuewen, et al.
Published: (2024)
by: Liu, Xuewen, et al.
Published: (2024)
Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge
by: Shen, Xuan, et al.
Published: (2023)
by: Shen, Xuan, et al.
Published: (2023)
ProtoQuant: Quantization of Prototypical Parts For General and Fine-Grained Image Classification
by: Janusz, Mikołaj, et al.
Published: (2026)
by: Janusz, Mikołaj, et al.
Published: (2026)
GSQ-Tuning: Group-Shared Exponents Integer in Fully Quantized Training for LLMs On-Device Fine-tuning
by: Zhou, Sifan, et al.
Published: (2025)
by: Zhou, Sifan, et al.
Published: (2025)
OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting
by: Hu, Xing, et al.
Published: (2025)
by: Hu, Xing, et al.
Published: (2025)
FlexQuant: Elastic Quantization Framework for Locally Hosted LLM on Edge Devices
by: Chai, Yuji, et al.
Published: (2025)
by: Chai, Yuji, et al.
Published: (2025)
ReSpinQuant: Efficient Layer-Wise LLM Quantization via Subspace Residual Rotation Approximation
by: Kim, Suyoung, et al.
Published: (2026)
by: Kim, Suyoung, et al.
Published: (2026)
KcMF: A Knowledge-compliant Framework for Schema and Entity Matching with Fine-tuning-free LLMs
by: Xu, Yongqin, et al.
Published: (2024)
by: Xu, Yongqin, et al.
Published: (2024)
ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly Transforms
by: Xu, Bingxin, et al.
Published: (2025)
by: Xu, Bingxin, et al.
Published: (2025)
Taming Sensitive Weights : Noise Perturbation Fine-tuning for Robust LLM Quantization
by: Wang, Dongwei, et al.
Published: (2024)
by: Wang, Dongwei, et al.
Published: (2024)
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
by: Yan, Xianglong, et al.
Published: (2026)
by: Yan, Xianglong, et al.
Published: (2026)
SplitQuantV2: Enhancing Low-Bit Quantization of LLMs Without GPUs
by: Song, Jaewoo, et al.
Published: (2025)
by: Song, Jaewoo, et al.
Published: (2025)
SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
by: Zhao, Zhixiong, et al.
Published: (2025)
by: Zhao, Zhixiong, et al.
Published: (2025)
You Only Fine-tune Once: Many-Shot In-Context Fine-Tuning for Large Language Models
by: He, Wenchong, et al.
Published: (2025)
by: He, Wenchong, et al.
Published: (2025)
MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM
by: Wang, Dongwei, et al.
Published: (2026)
by: Wang, Dongwei, et al.
Published: (2026)
HAFLQ: Heterogeneous Adaptive Federated LoRA Fine-tuned LLM with Quantization
by: Su, Yang, et al.
Published: (2024)
by: Su, Yang, et al.
Published: (2024)
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
by: Xiao, Guangxuan, et al.
Published: (2022)
by: Xiao, Guangxuan, et al.
Published: (2022)
PolarQuant: Quantizing KV Caches with Polar Transformation
by: Han, Insu, et al.
Published: (2025)
by: Han, Insu, et al.
Published: (2025)
AMAQ: Adaptive Mixed-bit Activation Quantization for Collaborative Parameter Efficient Fine-tuning
by: Song, Yurun, et al.
Published: (2025)
by: Song, Yurun, et al.
Published: (2025)
S$^{2}$FT: Efficient, Scalable and Generalizable LLM Fine-tuning by Structured Sparsity
by: Yang, Xinyu, et al.
Published: (2024)
by: Yang, Xinyu, et al.
Published: (2024)
LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models
by: Chen, Yukang, et al.
Published: (2023)
by: Chen, Yukang, et al.
Published: (2023)
ComplLLM: Fine-tuning LLMs to Discover Complementary Signals for Decision-making
by: Guo, Ziyang, et al.
Published: (2026)
by: Guo, Ziyang, et al.
Published: (2026)
RoLoRA: Fine-tuning Rotated Outlier-free LLMs for Effective Weight-Activation Quantization
by: Huang, Xijie, et al.
Published: (2024)
by: Huang, Xijie, et al.
Published: (2024)
Similar Items
-
EasyQuant: An Efficient Data-free Quantization Algorithm for LLMs
by: Tang, Hanlin, et al.
Published: (2024) -
Detect an Object At Once without Fine-tuning
by: Hao, Junyu, et al.
Published: (2024) -
InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
by: Li, Ke, et al.
Published: (2026) -
SliderQuant: Accurate Post-Training Quantization for LLMs
by: Wang, Shigeng, et al.
Published: (2026) -
Quantized Evolution Strategies: High-precision Fine-tuning of Quantized LLMs at Low-precision Cost
by: Xu, Yinggan, et al.
Published: (2026)