QuIP: 2-Bit Quantization of Large Language Models With Guarantees
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chee, Jerry, Cai, Yaohui, Kuleshov, Volodymyr, De Sa, Christopher |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
von: Tseng, Albert, et al.
Veröffentlicht: (2024)
von: Tseng, Albert, et al.
Veröffentlicht: (2024)
ModuLoRA: Finetuning 2-Bit LLMs on Consumer GPUs by Integrating with Modular Quantizers
von: Yin, Junjie, et al.
Veröffentlicht: (2023)
von: Yin, Junjie, et al.
Veröffentlicht: (2023)
ApiQ: Finetuning of 2-Bit Quantized Large Language Model
von: Liao, Baohao, et al.
Veröffentlicht: (2024)
von: Liao, Baohao, et al.
Veröffentlicht: (2024)
BAQ: Efficient Bit Allocation Quantization for Large Language Models
von: Zhang, Chao, et al.
Veröffentlicht: (2025)
von: Zhang, Chao, et al.
Veröffentlicht: (2025)
Active Preference Inference using Language Models and Probabilistic Reasoning
von: Piriyakulkij, Wasu Top, et al.
Veröffentlicht: (2023)
von: Piriyakulkij, Wasu Top, et al.
Veröffentlicht: (2023)
Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models
von: Park, Jungwoo, et al.
Veröffentlicht: (2025)
von: Park, Jungwoo, et al.
Veröffentlicht: (2025)
LCQ: Low-Rank Codebook based Quantization for Large Language Models
von: Cai, Wen-Pu, et al.
Veröffentlicht: (2024)
von: Cai, Wen-Pu, et al.
Veröffentlicht: (2024)
I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models
von: Hu, Xing, et al.
Veröffentlicht: (2024)
von: Hu, Xing, et al.
Veröffentlicht: (2024)
Multi-Bit Distortion-Free Watermarking for Large Language Models
von: Boroujeny, Massieh Kordi, et al.
Veröffentlicht: (2024)
von: Boroujeny, Massieh Kordi, et al.
Veröffentlicht: (2024)
QuAILoRA: Quantization-Aware Initialization for LoRA
von: Lawton, Neal, et al.
Veröffentlicht: (2024)
von: Lawton, Neal, et al.
Veröffentlicht: (2024)
Language Models with Conformal Factuality Guarantees
von: Mohri, Christopher, et al.
Veröffentlicht: (2024)
von: Mohri, Christopher, et al.
Veröffentlicht: (2024)
pQuant: Towards Effective Low-Bit Language Models via Decoupled Linear Quantization-Aware Training
von: Zhang, Wenzheng, et al.
Veröffentlicht: (2026)
von: Zhang, Wenzheng, et al.
Veröffentlicht: (2026)
FrameQuant: Flexible Low-Bit Quantization for Transformers
von: Adepu, Harshavardhan, et al.
Veröffentlicht: (2024)
von: Adepu, Harshavardhan, et al.
Veröffentlicht: (2024)
QuRating: Selecting High-Quality Data for Training Language Models
von: Wettig, Alexander, et al.
Veröffentlicht: (2024)
von: Wettig, Alexander, et al.
Veröffentlicht: (2024)
Simple and Effective Masked Diffusion Language Models
von: Sahoo, Subham Sekhar, et al.
Veröffentlicht: (2024)
von: Sahoo, Subham Sekhar, et al.
Veröffentlicht: (2024)
Quantization of Large Language Models with an Overdetermined Basis
von: Merkulov, Daniil, et al.
Veröffentlicht: (2024)
von: Merkulov, Daniil, et al.
Veröffentlicht: (2024)
LittleBit: Ultra Low-Bit Quantization via Latent Factorization
von: Lee, Banseok, et al.
Veröffentlicht: (2025)
von: Lee, Banseok, et al.
Veröffentlicht: (2025)
Diffusion Models With Learned Adaptive Noise
von: Sahoo, Subham Sekhar, et al.
Veröffentlicht: (2023)
von: Sahoo, Subham Sekhar, et al.
Veröffentlicht: (2023)
CBQ: Cross-Block Quantization for Large Language Models
von: Ding, Xin, et al.
Veröffentlicht: (2023)
von: Ding, Xin, et al.
Veröffentlicht: (2023)
FBQuant: FeedBack Quantization for Large Language Models
von: Liu, Yijiang, et al.
Veröffentlicht: (2025)
von: Liu, Yijiang, et al.
Veröffentlicht: (2025)
How Quantization Shapes Bias in Large Language Models
von: Marcuzzi, Federico, et al.
Veröffentlicht: (2025)
von: Marcuzzi, Federico, et al.
Veröffentlicht: (2025)
Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens
von: Ouyang, Xu, et al.
Veröffentlicht: (2024)
von: Ouyang, Xu, et al.
Veröffentlicht: (2024)
The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
von: Ma, Shuming, et al.
Veröffentlicht: (2024)
von: Ma, Shuming, et al.
Veröffentlicht: (2024)
QLLM: Accurate and Efficient Low-Bitwidth Quantization for Large Language Models
von: Liu, Jing, et al.
Veröffentlicht: (2023)
von: Liu, Jing, et al.
Veröffentlicht: (2023)
OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models
von: Shao, Wenqi, et al.
Veröffentlicht: (2023)
von: Shao, Wenqi, et al.
Veröffentlicht: (2023)
Scaling Laws for Post Training Quantized Large Language Models
von: Xu, Zifei, et al.
Veröffentlicht: (2024)
von: Xu, Zifei, et al.
Veröffentlicht: (2024)
Extreme Compression of Large Language Models via Additive Quantization
von: Egiazarian, Vage, et al.
Veröffentlicht: (2024)
von: Egiazarian, Vage, et al.
Veröffentlicht: (2024)
TFL: Targeted Bit-Flip Attack on Large Language Model
von: Guo, Jingkai, et al.
Veröffentlicht: (2026)
von: Guo, Jingkai, et al.
Veröffentlicht: (2026)
DUEL: Exact Likelihood for Masked Diffusion via Deterministic Unmasking
von: Turok, Gilad, et al.
Veröffentlicht: (2026)
von: Turok, Gilad, et al.
Veröffentlicht: (2026)
On the Compressibility of Quantized Large Language Models
von: Mao, Yu, et al.
Veröffentlicht: (2024)
von: Mao, Yu, et al.
Veröffentlicht: (2024)
RSAVQ: Riemannian Sensitivity-Aware Vector Quantization for Large Language Models
von: Xu, Zukang, et al.
Veröffentlicht: (2025)
von: Xu, Zukang, et al.
Veröffentlicht: (2025)
Optimizing Large Language Model Training Using FP4 Quantization
von: Wang, Ruizhe, et al.
Veröffentlicht: (2025)
von: Wang, Ruizhe, et al.
Veröffentlicht: (2025)
Foundations of Large Language Model Compression -- Part 1: Weight Quantization
von: Young, Sean I.
Veröffentlicht: (2024)
von: Young, Sean I.
Veröffentlicht: (2024)
BitDelta: Your Fine-Tune May Only Be Worth One Bit
von: Liu, James, et al.
Veröffentlicht: (2024)
von: Liu, James, et al.
Veröffentlicht: (2024)
H1B-KV: Hybrid One-Bit Caches for Memory-Efficient Large Language Model Inference
von: Vejendla, Harshil
Veröffentlicht: (2025)
von: Vejendla, Harshil
Veröffentlicht: (2025)
SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models
von: Duanmu, Haojie, et al.
Veröffentlicht: (2024)
von: Duanmu, Haojie, et al.
Veröffentlicht: (2024)
PrefixQuant: Eliminating Outliers by Prefixed Tokens for Large Language Models Quantization
von: Chen, Mengzhao, et al.
Veröffentlicht: (2024)
von: Chen, Mengzhao, et al.
Veröffentlicht: (2024)
When are 1.58 bits enough? A Bottom-up Exploration of BitNet Quantization
von: Nielsen, Jacob, et al.
Veröffentlicht: (2024)
von: Nielsen, Jacob, et al.
Veröffentlicht: (2024)
SBFA: Single Sneaky Bit Flip Attack to Break Large Language Models
von: Guo, Jingkai, et al.
Veröffentlicht: (2025)
von: Guo, Jingkai, et al.
Veröffentlicht: (2025)
Technical Report: Activation Residual Hessian Quantization (ARHQ) for Low-Bit LLM Quantization
von: Wang, YiFeng, et al.
Veröffentlicht: (2026)
von: Wang, YiFeng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
von: Tseng, Albert, et al.
Veröffentlicht: (2024) -
ModuLoRA: Finetuning 2-Bit LLMs on Consumer GPUs by Integrating with Modular Quantizers
von: Yin, Junjie, et al.
Veröffentlicht: (2023) -
ApiQ: Finetuning of 2-Bit Quantized Large Language Model
von: Liao, Baohao, et al.
Veröffentlicht: (2024) -
BAQ: Efficient Bit Allocation Quantization for Large Language Models
von: Zhang, Chao, et al.
Veröffentlicht: (2025) -
Active Preference Inference using Language Models and Probabilistic Reasoning
von: Piriyakulkij, Wasu Top, et al.
Veröffentlicht: (2023)