One QuantLLM for ALL: Fine-tuning Quantized LLMs Once for Efficient Deployments
Fuente:
arXiv
Salvato in:
| Autori principali: | Yi, Ke, Xu, Yuhui, Chang, Heng, Tang, Chen, Meng, Yuan, Zhang, Tong, Li, Jia |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
EasyQuant: An Efficient Data-free Quantization Algorithm for LLMs
di: Tang, Hanlin, et al.
Pubblicazione: (2024)
di: Tang, Hanlin, et al.
Pubblicazione: (2024)
Detect an Object At Once without Fine-tuning
di: Hao, Junyu, et al.
Pubblicazione: (2024)
di: Hao, Junyu, et al.
Pubblicazione: (2024)
InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
di: Li, Ke, et al.
Pubblicazione: (2026)
di: Li, Ke, et al.
Pubblicazione: (2026)
SliderQuant: Accurate Post-Training Quantization for LLMs
di: Wang, Shigeng, et al.
Pubblicazione: (2026)
di: Wang, Shigeng, et al.
Pubblicazione: (2026)
Quantized Evolution Strategies: High-precision Fine-tuning of Quantized LLMs at Low-precision Cost
di: Xu, Yinggan, et al.
Pubblicazione: (2026)
di: Xu, Yinggan, et al.
Pubblicazione: (2026)
PrivQuant: Communication-Efficient Private Inference with Quantized Network/Protocol Co-Optimization
di: Xu, Tianshi, et al.
Pubblicazione: (2024)
di: Xu, Tianshi, et al.
Pubblicazione: (2024)
Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
di: Zhang, Tuo, et al.
Pubblicazione: (2025)
di: Zhang, Tuo, et al.
Pubblicazione: (2025)
Quant.npu: Enabling Efficient Mobile NPU Inference for on-device LLMs via Fully Static Quantization
di: Zhang, Jinghe, et al.
Pubblicazione: (2026)
di: Zhang, Jinghe, et al.
Pubblicazione: (2026)
DuQuant++: Fine-grained Rotation Enhances Microscaling FP4 Quantization
di: Lin, Haokun, et al.
Pubblicazione: (2026)
di: Lin, Haokun, et al.
Pubblicazione: (2026)
DFQ-ViT: Data-Free Quantization for Vision Transformers without Fine-tuning
di: Tong, Yujia, et al.
Pubblicazione: (2025)
di: Tong, Yujia, et al.
Pubblicazione: (2025)
NestQuant: Nested Lattice Quantization for Matrix Products and LLMs
di: Savkin, Semyon, et al.
Pubblicazione: (2025)
di: Savkin, Semyon, et al.
Pubblicazione: (2025)
Q-realign: Piggybacking Realignment on Quantization for Safe and Efficient LLM Deployment
di: Tan, Qitao, et al.
Pubblicazione: (2026)
di: Tan, Qitao, et al.
Pubblicazione: (2026)
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
di: Xu, Zukang, et al.
Pubblicazione: (2025)
di: Xu, Zukang, et al.
Pubblicazione: (2025)
On-the-Fly Adaptation to Quantization: Configuration-Aware LoRA for Efficient Fine-Tuning of Quantized LLMs
di: Ye, Rongguang, et al.
Pubblicazione: (2025)
di: Ye, Rongguang, et al.
Pubblicazione: (2025)
QuantLRM: Quantization of Large Reasoning Models via Fine-Tuning Signals
di: Zhang, Nan, et al.
Pubblicazione: (2026)
di: Zhang, Nan, et al.
Pubblicazione: (2026)
ALL-FEM: Agentic Large Language models Fine-tuned for Finite Element Methods
di: Deotale, Rushikesh, et al.
Pubblicazione: (2026)
di: Deotale, Rushikesh, et al.
Pubblicazione: (2026)
ZorBA: Zeroth-order Federated Fine-tuning of LLMs with Heterogeneous Block Activation
di: Meng, Chuiyang, et al.
Pubblicazione: (2026)
di: Meng, Chuiyang, et al.
Pubblicazione: (2026)
DilateQuant: Accurate and Efficient Diffusion Quantization via Weight Dilation
di: Liu, Xuewen, et al.
Pubblicazione: (2024)
di: Liu, Xuewen, et al.
Pubblicazione: (2024)
Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge
di: Shen, Xuan, et al.
Pubblicazione: (2023)
di: Shen, Xuan, et al.
Pubblicazione: (2023)
ProtoQuant: Quantization of Prototypical Parts For General and Fine-Grained Image Classification
di: Janusz, Mikołaj, et al.
Pubblicazione: (2026)
di: Janusz, Mikołaj, et al.
Pubblicazione: (2026)
GSQ-Tuning: Group-Shared Exponents Integer in Fully Quantized Training for LLMs On-Device Fine-tuning
di: Zhou, Sifan, et al.
Pubblicazione: (2025)
di: Zhou, Sifan, et al.
Pubblicazione: (2025)
OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting
di: Hu, Xing, et al.
Pubblicazione: (2025)
di: Hu, Xing, et al.
Pubblicazione: (2025)
FlexQuant: Elastic Quantization Framework for Locally Hosted LLM on Edge Devices
di: Chai, Yuji, et al.
Pubblicazione: (2025)
di: Chai, Yuji, et al.
Pubblicazione: (2025)
ReSpinQuant: Efficient Layer-Wise LLM Quantization via Subspace Residual Rotation Approximation
di: Kim, Suyoung, et al.
Pubblicazione: (2026)
di: Kim, Suyoung, et al.
Pubblicazione: (2026)
KcMF: A Knowledge-compliant Framework for Schema and Entity Matching with Fine-tuning-free LLMs
di: Xu, Yongqin, et al.
Pubblicazione: (2024)
di: Xu, Yongqin, et al.
Pubblicazione: (2024)
ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly Transforms
di: Xu, Bingxin, et al.
Pubblicazione: (2025)
di: Xu, Bingxin, et al.
Pubblicazione: (2025)
Taming Sensitive Weights : Noise Perturbation Fine-tuning for Robust LLM Quantization
di: Wang, Dongwei, et al.
Pubblicazione: (2024)
di: Wang, Dongwei, et al.
Pubblicazione: (2024)
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
di: Yan, Xianglong, et al.
Pubblicazione: (2026)
di: Yan, Xianglong, et al.
Pubblicazione: (2026)
SplitQuantV2: Enhancing Low-Bit Quantization of LLMs Without GPUs
di: Song, Jaewoo, et al.
Pubblicazione: (2025)
di: Song, Jaewoo, et al.
Pubblicazione: (2025)
SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
di: Zhao, Zhixiong, et al.
Pubblicazione: (2025)
di: Zhao, Zhixiong, et al.
Pubblicazione: (2025)
You Only Fine-tune Once: Many-Shot In-Context Fine-Tuning for Large Language Models
di: He, Wenchong, et al.
Pubblicazione: (2025)
di: He, Wenchong, et al.
Pubblicazione: (2025)
MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM
di: Wang, Dongwei, et al.
Pubblicazione: (2026)
di: Wang, Dongwei, et al.
Pubblicazione: (2026)
HAFLQ: Heterogeneous Adaptive Federated LoRA Fine-tuned LLM with Quantization
di: Su, Yang, et al.
Pubblicazione: (2024)
di: Su, Yang, et al.
Pubblicazione: (2024)
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
di: Xiao, Guangxuan, et al.
Pubblicazione: (2022)
di: Xiao, Guangxuan, et al.
Pubblicazione: (2022)
PolarQuant: Quantizing KV Caches with Polar Transformation
di: Han, Insu, et al.
Pubblicazione: (2025)
di: Han, Insu, et al.
Pubblicazione: (2025)
AMAQ: Adaptive Mixed-bit Activation Quantization for Collaborative Parameter Efficient Fine-tuning
di: Song, Yurun, et al.
Pubblicazione: (2025)
di: Song, Yurun, et al.
Pubblicazione: (2025)
S$^{2}$FT: Efficient, Scalable and Generalizable LLM Fine-tuning by Structured Sparsity
di: Yang, Xinyu, et al.
Pubblicazione: (2024)
di: Yang, Xinyu, et al.
Pubblicazione: (2024)
LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models
di: Chen, Yukang, et al.
Pubblicazione: (2023)
di: Chen, Yukang, et al.
Pubblicazione: (2023)
ComplLLM: Fine-tuning LLMs to Discover Complementary Signals for Decision-making
di: Guo, Ziyang, et al.
Pubblicazione: (2026)
di: Guo, Ziyang, et al.
Pubblicazione: (2026)
RoLoRA: Fine-tuning Rotated Outlier-free LLMs for Effective Weight-Activation Quantization
di: Huang, Xijie, et al.
Pubblicazione: (2024)
di: Huang, Xijie, et al.
Pubblicazione: (2024)
Documenti analoghi
-
EasyQuant: An Efficient Data-free Quantization Algorithm for LLMs
di: Tang, Hanlin, et al.
Pubblicazione: (2024) -
Detect an Object At Once without Fine-tuning
di: Hao, Junyu, et al.
Pubblicazione: (2024) -
InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
di: Li, Ke, et al.
Pubblicazione: (2026) -
SliderQuant: Accurate Post-Training Quantization for LLMs
di: Wang, Shigeng, et al.
Pubblicazione: (2026) -
Quantized Evolution Strategies: High-precision Fine-tuning of Quantized LLMs at Low-precision Cost
di: Xu, Yinggan, et al.
Pubblicazione: (2026)