QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tseng, Albert, Chee, Jerry, Sun, Qingyao, Kuleshov, Volodymyr, De Sa, Christopher |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
QuIP: 2-Bit Quantization of Large Language Models With Guarantees
von: Chee, Jerry, et al.
Veröffentlicht: (2023)
von: Chee, Jerry, et al.
Veröffentlicht: (2023)
QTIP: Quantization with Trellises and Incoherence Processing
von: Tseng, Albert, et al.
Veröffentlicht: (2024)
von: Tseng, Albert, et al.
Veröffentlicht: (2024)
ModuLoRA: Finetuning 2-Bit LLMs on Consumer GPUs by Integrating with Modular Quantizers
von: Yin, Junjie, et al.
Veröffentlicht: (2023)
von: Yin, Junjie, et al.
Veröffentlicht: (2023)
Locally Coherent, Globally Incoherent: Bounding Compositional Incoherence in Multi-Component LLM Agents
von: Kotawala, Anany
Veröffentlicht: (2026)
von: Kotawala, Anany
Veröffentlicht: (2026)
Active Preference Inference using Language Models and Probabilistic Reasoning
von: Piriyakulkij, Wasu Top, et al.
Veröffentlicht: (2023)
von: Piriyakulkij, Wasu Top, et al.
Veröffentlicht: (2023)
L$^3$: Large Lookup Layers
von: Tseng, Albert, et al.
Veröffentlicht: (2026)
von: Tseng, Albert, et al.
Veröffentlicht: (2026)
Model-Preserving Adaptive Rounding
von: Tseng, Albert, et al.
Veröffentlicht: (2025)
von: Tseng, Albert, et al.
Veröffentlicht: (2025)
Calibrating Beyond English: Language Diversity for Better Quantized Multilingual LLM
von: Chimoto, Everlyn Asiko, et al.
Veröffentlicht: (2026)
von: Chimoto, Everlyn Asiko, et al.
Veröffentlicht: (2026)
Incoherent Probability Judgments in Large Language Models
von: Zhu, Jian-Qiao, et al.
Veröffentlicht: (2024)
von: Zhu, Jian-Qiao, et al.
Veröffentlicht: (2024)
The Diffusion Duality
von: Sahoo, Subham Sekhar, et al.
Veröffentlicht: (2025)
von: Sahoo, Subham Sekhar, et al.
Veröffentlicht: (2025)
Incoherence as Oracle-less Measure of Error in LLM-Based Code Generation
von: Valentin, Thomas, et al.
Veröffentlicht: (2025)
von: Valentin, Thomas, et al.
Veröffentlicht: (2025)
Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation
von: Manvi, Rohin, et al.
Veröffentlicht: (2024)
von: Manvi, Rohin, et al.
Veröffentlicht: (2024)
QuIP: A P4 Quantum Internet Protocol Prototyping Framework
von: Kozlowski, Wojciech, et al.
Veröffentlicht: (2024)
von: Kozlowski, Wojciech, et al.
Veröffentlicht: (2024)
Codebook-Injected Dialogue Segmentation for Multi-Utterance Constructs Annotation: LLM-Assisted and Gold-Label-Free Evaluation
von: Lee, Jinsook, et al.
Veröffentlicht: (2026)
von: Lee, Jinsook, et al.
Veröffentlicht: (2026)
Tokenizer Fertility and Zero-Shot Performance of Foundation Models on Ukrainian Legal Text: A Comparative Study
von: Ovcharov, Volodymyr
Veröffentlicht: (2026)
von: Ovcharov, Volodymyr
Veröffentlicht: (2026)
SwaQuAD-24: QA Benchmark Dataset in Swahili
von: Kondoro, Alfred Malengo
Veröffentlicht: (2024)
von: Kondoro, Alfred Malengo
Veröffentlicht: (2024)
Cross-Lingual Jailbreak Detection via Semantic Codebooks
von: Alanova, Shirin, et al.
Veröffentlicht: (2026)
von: Alanova, Shirin, et al.
Veröffentlicht: (2026)
ACCEPT: Adaptive Codebook for Composite and Efficient Prompt Tuning
von: Lin, Yu-Chen, et al.
Veröffentlicht: (2024)
von: Lin, Yu-Chen, et al.
Veröffentlicht: (2024)
Quantization Meets Reasoning: Exploring LLM Low-Bit Quantization Degradation for Mathematical Reasoning
von: Li, Zhen, et al.
Veröffentlicht: (2025)
von: Li, Zhen, et al.
Veröffentlicht: (2025)
Catastrophic Failure of LLM Unlearning via Quantization
von: Zhang, Zhiwei, et al.
Veröffentlicht: (2024)
von: Zhang, Zhiwei, et al.
Veröffentlicht: (2024)
Executable Code Actions Elicit Better LLM Agents
von: Wang, Xingyao, et al.
Veröffentlicht: (2024)
von: Wang, Xingyao, et al.
Veröffentlicht: (2024)
Learning From Mistakes Makes LLM Better Reasoner
von: An, Shengnan, et al.
Veröffentlicht: (2023)
von: An, Shengnan, et al.
Veröffentlicht: (2023)
IP Leakage Attacks Targeting LLM-Based Multi-Agent Systems
von: Wang, Liwen, et al.
Veröffentlicht: (2025)
von: Wang, Liwen, et al.
Veröffentlicht: (2025)
RSQ: Learning from Important Tokens Leads to Better Quantized LLMs
von: Sung, Yi-Lin, et al.
Veröffentlicht: (2025)
von: Sung, Yi-Lin, et al.
Veröffentlicht: (2025)
Simple and Effective Masked Diffusion Language Models
von: Sahoo, Subham Sekhar, et al.
Veröffentlicht: (2024)
von: Sahoo, Subham Sekhar, et al.
Veröffentlicht: (2024)
BanglaQuAD: A Bengali Open-domain Question Answering Dataset
von: Rony, Md Rashad Al Hasan, et al.
Veröffentlicht: (2024)
von: Rony, Md Rashad Al Hasan, et al.
Veröffentlicht: (2024)
Highly Optimized Kernels and Fine-Grained Codebooks for LLM Inference on Arm CPUs
von: Gope, Dibakar, et al.
Veröffentlicht: (2024)
von: Gope, Dibakar, et al.
Veröffentlicht: (2024)
The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language Models
von: Chen, Yanjun, et al.
Veröffentlicht: (2024)
von: Chen, Yanjun, et al.
Veröffentlicht: (2024)
Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models
von: Liu, Ruikang, et al.
Veröffentlicht: (2025)
von: Liu, Ruikang, et al.
Veröffentlicht: (2025)
Blending Is All You Need: Cheaper, Better Alternative to Trillion-Parameters LLM
von: Lu, Xiaoding, et al.
Veröffentlicht: (2024)
von: Lu, Xiaoding, et al.
Veröffentlicht: (2024)
Turning LLM Activations Quantization-Friendly
von: Czakó, Patrik, et al.
Veröffentlicht: (2025)
von: Czakó, Patrik, et al.
Veröffentlicht: (2025)
Better LLM Reasoning via Dual-Play
von: Zhang, Zhengxin, et al.
Veröffentlicht: (2025)
von: Zhang, Zhengxin, et al.
Veröffentlicht: (2025)
Improve Large Language Model Systems with User Logs
von: Wang, Changyue, et al.
Veröffentlicht: (2026)
von: Wang, Changyue, et al.
Veröffentlicht: (2026)
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
von: Lee, Jaehyeok, et al.
Veröffentlicht: (2026)
von: Lee, Jaehyeok, et al.
Veröffentlicht: (2026)
Codebook Reduction and Saturation: Novel observations on Inductive Thematic Saturation for Large Language Models and initial coding in Thematic Analysis
von: De Paoli, Stefano, et al.
Veröffentlicht: (2025)
von: De Paoli, Stefano, et al.
Veröffentlicht: (2025)
Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
Do Before You Judge: Self-Reference as a Pathway to Better LLM Evaluation
von: Lin, Wei-Hsiang, et al.
Veröffentlicht: (2025)
von: Lin, Wei-Hsiang, et al.
Veröffentlicht: (2025)
Training With "Paraphrasing the Original Text" Teaches LLM to Better Retrieve in Long-context Tasks
von: Yu, Yijiong, et al.
Veröffentlicht: (2023)
von: Yu, Yijiong, et al.
Veröffentlicht: (2023)
AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
HARP: Hadamard-Preconditioned Adaptive Rotation Processor for Extreme LLM Quantization
von: Zagitov, Artur, et al.
Veröffentlicht: (2026)
von: Zagitov, Artur, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
QuIP: 2-Bit Quantization of Large Language Models With Guarantees
von: Chee, Jerry, et al.
Veröffentlicht: (2023) -
QTIP: Quantization with Trellises and Incoherence Processing
von: Tseng, Albert, et al.
Veröffentlicht: (2024) -
ModuLoRA: Finetuning 2-Bit LLMs on Consumer GPUs by Integrating with Modular Quantizers
von: Yin, Junjie, et al.
Veröffentlicht: (2023) -
Locally Coherent, Globally Incoherent: Bounding Compositional Incoherence in Multi-Component LLM Agents
von: Kotawala, Anany
Veröffentlicht: (2026) -
Active Preference Inference using Language Models and Probabilistic Reasoning
von: Piriyakulkij, Wasu Top, et al.
Veröffentlicht: (2023)