QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Fuente:
arXiv
Saved in:
| Main Authors: | Tseng, Albert, Chee, Jerry, Sun, Qingyao, Kuleshov, Volodymyr, De Sa, Christopher |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
QuIP: 2-Bit Quantization of Large Language Models With Guarantees
by: Chee, Jerry, et al.
Published: (2023)
by: Chee, Jerry, et al.
Published: (2023)
QTIP: Quantization with Trellises and Incoherence Processing
by: Tseng, Albert, et al.
Published: (2024)
by: Tseng, Albert, et al.
Published: (2024)
ModuLoRA: Finetuning 2-Bit LLMs on Consumer GPUs by Integrating with Modular Quantizers
by: Yin, Junjie, et al.
Published: (2023)
by: Yin, Junjie, et al.
Published: (2023)
Locally Coherent, Globally Incoherent: Bounding Compositional Incoherence in Multi-Component LLM Agents
by: Kotawala, Anany
Published: (2026)
by: Kotawala, Anany
Published: (2026)
Active Preference Inference using Language Models and Probabilistic Reasoning
by: Piriyakulkij, Wasu Top, et al.
Published: (2023)
by: Piriyakulkij, Wasu Top, et al.
Published: (2023)
L$^3$: Large Lookup Layers
by: Tseng, Albert, et al.
Published: (2026)
by: Tseng, Albert, et al.
Published: (2026)
Model-Preserving Adaptive Rounding
by: Tseng, Albert, et al.
Published: (2025)
by: Tseng, Albert, et al.
Published: (2025)
Calibrating Beyond English: Language Diversity for Better Quantized Multilingual LLM
by: Chimoto, Everlyn Asiko, et al.
Published: (2026)
by: Chimoto, Everlyn Asiko, et al.
Published: (2026)
Incoherent Probability Judgments in Large Language Models
by: Zhu, Jian-Qiao, et al.
Published: (2024)
by: Zhu, Jian-Qiao, et al.
Published: (2024)
The Diffusion Duality
by: Sahoo, Subham Sekhar, et al.
Published: (2025)
by: Sahoo, Subham Sekhar, et al.
Published: (2025)
Incoherence as Oracle-less Measure of Error in LLM-Based Code Generation
by: Valentin, Thomas, et al.
Published: (2025)
by: Valentin, Thomas, et al.
Published: (2025)
Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation
by: Manvi, Rohin, et al.
Published: (2024)
by: Manvi, Rohin, et al.
Published: (2024)
QuIP: A P4 Quantum Internet Protocol Prototyping Framework
by: Kozlowski, Wojciech, et al.
Published: (2024)
by: Kozlowski, Wojciech, et al.
Published: (2024)
Codebook-Injected Dialogue Segmentation for Multi-Utterance Constructs Annotation: LLM-Assisted and Gold-Label-Free Evaluation
by: Lee, Jinsook, et al.
Published: (2026)
by: Lee, Jinsook, et al.
Published: (2026)
Tokenizer Fertility and Zero-Shot Performance of Foundation Models on Ukrainian Legal Text: A Comparative Study
by: Ovcharov, Volodymyr
Published: (2026)
by: Ovcharov, Volodymyr
Published: (2026)
SwaQuAD-24: QA Benchmark Dataset in Swahili
by: Kondoro, Alfred Malengo
Published: (2024)
by: Kondoro, Alfred Malengo
Published: (2024)
Cross-Lingual Jailbreak Detection via Semantic Codebooks
by: Alanova, Shirin, et al.
Published: (2026)
by: Alanova, Shirin, et al.
Published: (2026)
ACCEPT: Adaptive Codebook for Composite and Efficient Prompt Tuning
by: Lin, Yu-Chen, et al.
Published: (2024)
by: Lin, Yu-Chen, et al.
Published: (2024)
Quantization Meets Reasoning: Exploring LLM Low-Bit Quantization Degradation for Mathematical Reasoning
by: Li, Zhen, et al.
Published: (2025)
by: Li, Zhen, et al.
Published: (2025)
Catastrophic Failure of LLM Unlearning via Quantization
by: Zhang, Zhiwei, et al.
Published: (2024)
by: Zhang, Zhiwei, et al.
Published: (2024)
Executable Code Actions Elicit Better LLM Agents
by: Wang, Xingyao, et al.
Published: (2024)
by: Wang, Xingyao, et al.
Published: (2024)
Learning From Mistakes Makes LLM Better Reasoner
by: An, Shengnan, et al.
Published: (2023)
by: An, Shengnan, et al.
Published: (2023)
IP Leakage Attacks Targeting LLM-Based Multi-Agent Systems
by: Wang, Liwen, et al.
Published: (2025)
by: Wang, Liwen, et al.
Published: (2025)
RSQ: Learning from Important Tokens Leads to Better Quantized LLMs
by: Sung, Yi-Lin, et al.
Published: (2025)
by: Sung, Yi-Lin, et al.
Published: (2025)
Simple and Effective Masked Diffusion Language Models
by: Sahoo, Subham Sekhar, et al.
Published: (2024)
by: Sahoo, Subham Sekhar, et al.
Published: (2024)
BanglaQuAD: A Bengali Open-domain Question Answering Dataset
by: Rony, Md Rashad Al Hasan, et al.
Published: (2024)
by: Rony, Md Rashad Al Hasan, et al.
Published: (2024)
Highly Optimized Kernels and Fine-Grained Codebooks for LLM Inference on Arm CPUs
by: Gope, Dibakar, et al.
Published: (2024)
by: Gope, Dibakar, et al.
Published: (2024)
The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language Models
by: Chen, Yanjun, et al.
Published: (2024)
by: Chen, Yanjun, et al.
Published: (2024)
Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models
by: Liu, Ruikang, et al.
Published: (2025)
by: Liu, Ruikang, et al.
Published: (2025)
Blending Is All You Need: Cheaper, Better Alternative to Trillion-Parameters LLM
by: Lu, Xiaoding, et al.
Published: (2024)
by: Lu, Xiaoding, et al.
Published: (2024)
Turning LLM Activations Quantization-Friendly
by: Czakó, Patrik, et al.
Published: (2025)
by: Czakó, Patrik, et al.
Published: (2025)
Better LLM Reasoning via Dual-Play
by: Zhang, Zhengxin, et al.
Published: (2025)
by: Zhang, Zhengxin, et al.
Published: (2025)
Improve Large Language Model Systems with User Logs
by: Wang, Changyue, et al.
Published: (2026)
by: Wang, Changyue, et al.
Published: (2026)
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
by: Lee, Jaehyeok, et al.
Published: (2026)
by: Lee, Jaehyeok, et al.
Published: (2026)
Codebook Reduction and Saturation: Novel observations on Inductive Thematic Saturation for Large Language Models and initial coding in Thematic Analysis
by: De Paoli, Stefano, et al.
Published: (2025)
by: De Paoli, Stefano, et al.
Published: (2025)
Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling
by: Bansal, Hritik, et al.
Published: (2024)
by: Bansal, Hritik, et al.
Published: (2024)
Do Before You Judge: Self-Reference as a Pathway to Better LLM Evaluation
by: Lin, Wei-Hsiang, et al.
Published: (2025)
by: Lin, Wei-Hsiang, et al.
Published: (2025)
Training With "Paraphrasing the Original Text" Teaches LLM to Better Retrieve in Long-context Tasks
by: Yu, Yijiong, et al.
Published: (2023)
by: Yu, Yijiong, et al.
Published: (2023)
AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
HARP: Hadamard-Preconditioned Adaptive Rotation Processor for Extreme LLM Quantization
by: Zagitov, Artur, et al.
Published: (2026)
by: Zagitov, Artur, et al.
Published: (2026)
Similar Items
-
QuIP: 2-Bit Quantization of Large Language Models With Guarantees
by: Chee, Jerry, et al.
Published: (2023) -
QTIP: Quantization with Trellises and Incoherence Processing
by: Tseng, Albert, et al.
Published: (2024) -
ModuLoRA: Finetuning 2-Bit LLMs on Consumer GPUs by Integrating with Modular Quantizers
by: Yin, Junjie, et al.
Published: (2023) -
Locally Coherent, Globally Incoherent: Bounding Compositional Incoherence in Multi-Component LLM Agents
by: Kotawala, Anany
Published: (2026) -
Active Preference Inference using Language Models and Probabilistic Reasoning
by: Piriyakulkij, Wasu Top, et al.
Published: (2023)