Which Quantization Should I Use? A Unified Evaluation of llama.cpp Quantization on Llama-3.1-8B-Instruct
Fuente:
arXiv
Salvato in:
| Autore principale: | Kurt, Uygar |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Forensic Implications of Localized AI: Artifact Analysis of Ollama, LM Studio, and llama.cpp
di: Murtuza, Shariq
Pubblicazione: (2026)
di: Murtuza, Shariq
Pubblicazione: (2026)
Pruning vs Quantization: Which is Better?
di: Kuzmin, Andrey, et al.
Pubblicazione: (2023)
di: Kuzmin, Andrey, et al.
Pubblicazione: (2023)
Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders
di: He, Zhengfu, et al.
Pubblicazione: (2024)
di: He, Zhengfu, et al.
Pubblicazione: (2024)
Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct
di: Ackerman, Christopher, et al.
Pubblicazione: (2024)
di: Ackerman, Christopher, et al.
Pubblicazione: (2024)
Domain Adaptation of Llama3-70B-Instruct through Continual Pre-Training and Model Merging: A Comprehensive Evaluation
di: Siriwardhana, Shamane, et al.
Pubblicazione: (2024)
di: Siriwardhana, Shamane, et al.
Pubblicazione: (2024)
Llama-3.1-FoundationAI-SecurityLLM-Reasoning-8B Technical Report
di: Yang, Zhuoran, et al.
Pubblicazione: (2026)
di: Yang, Zhuoran, et al.
Pubblicazione: (2026)
Production-Grade Local LLM Inference on Apple Silicon: A Comparative Study of MLX, MLC-LLM, Ollama, llama.cpp, and PyTorch MPS
di: Rajesh, Varun, et al.
Pubblicazione: (2025)
di: Rajesh, Varun, et al.
Pubblicazione: (2025)
FP8 Quantization: The Power of the Exponent
di: Kuzmin, Andrey, et al.
Pubblicazione: (2022)
di: Kuzmin, Andrey, et al.
Pubblicazione: (2022)
Compression Scaling Laws:Unifying Sparsity and Quantization
di: Frantar, Elias, et al.
Pubblicazione: (2025)
di: Frantar, Elias, et al.
Pubblicazione: (2025)
UniComp: A Unified Evaluation of Large Language Model Compression via Pruning, Quantization and Distillation
di: von Rad, Jonathan, et al.
Pubblicazione: (2026)
di: von Rad, Jonathan, et al.
Pubblicazione: (2026)
A Quantized VAE-MLP Botnet Detection Model: A Systematic Evaluation of Quantization-Aware Training and Post-Training Quantization Strategies
di: Wasswa, Hassan, et al.
Pubblicazione: (2025)
di: Wasswa, Hassan, et al.
Pubblicazione: (2025)
QuantKAN: A Unified Quantization Framework for Kolmogorov Arnold Networks
di: Fuad, Kazi Ahmed Asif, et al.
Pubblicazione: (2025)
di: Fuad, Kazi Ahmed Asif, et al.
Pubblicazione: (2025)
Which Augmentation Should I Use? An Empirical Investigation of Augmentations for Self-Supervised Phonocardiogram Representation Learning
di: Ballas, Aristotelis, et al.
Pubblicazione: (2023)
di: Ballas, Aristotelis, et al.
Pubblicazione: (2023)
SqueezeLLM: Dense-and-Sparse Quantization
di: Kim, Sehoon, et al.
Pubblicazione: (2023)
di: Kim, Sehoon, et al.
Pubblicazione: (2023)
Benchmarking Post-Training Quantization in LLMs: Comprehensive Taxonomy, Unified Evaluation, and Comparative Analysis
di: Zhao, Jiaqi, et al.
Pubblicazione: (2025)
di: Zhao, Jiaqi, et al.
Pubblicazione: (2025)
Unified Stochastic Framework for Neural Network Quantization and Pruning
di: Zhang, Haoyu, et al.
Pubblicazione: (2024)
di: Zhang, Haoyu, et al.
Pubblicazione: (2024)
An Empirical Study of Qwen3 Quantization
di: Zheng, Xingyu, et al.
Pubblicazione: (2025)
di: Zheng, Xingyu, et al.
Pubblicazione: (2025)
A Comprehensive Evaluation on Quantization Techniques for Large Language Models
di: Liu, Yutong, et al.
Pubblicazione: (2025)
di: Liu, Yutong, et al.
Pubblicazione: (2025)
A Systematic Evaluation of On-Device LLMs: Quantization, Performance, and Resources
di: Song, Qingyu, et al.
Pubblicazione: (2025)
di: Song, Qingyu, et al.
Pubblicazione: (2025)
Understanding Efficiency: Quantization, Batching, and Serving Strategies in LLM Energy Use
di: Delavande, Julien, et al.
Pubblicazione: (2026)
di: Delavande, Julien, et al.
Pubblicazione: (2026)
Activation Sensitivity as a Unifying Principle for Post-Training Quantization
di: Xu, Bruce Changlong
Pubblicazione: (2026)
di: Xu, Bruce Changlong
Pubblicazione: (2026)
QFT: Quantized Full-parameter Tuning of LLMs with Affordable Resources
di: Li, Zhikai, et al.
Pubblicazione: (2023)
di: Li, Zhikai, et al.
Pubblicazione: (2023)
SKIM: Any-bit Quantization Pushing The Limits of Post-Training Quantization
di: Bai, Runsheng, et al.
Pubblicazione: (2024)
di: Bai, Runsheng, et al.
Pubblicazione: (2024)
Learning Unified User Quantized Tokenizers for User Representation
di: He, Chuan, et al.
Pubblicazione: (2025)
di: He, Chuan, et al.
Pubblicazione: (2025)
Layer-wise Quantization for Quantized Optimistic Dual Averaging
di: Nguyen, Anh Duc, et al.
Pubblicazione: (2025)
di: Nguyen, Anh Duc, et al.
Pubblicazione: (2025)
Operationalizing Quantized Disentanglement
di: Barin-Pacela, Vitoria, et al.
Pubblicazione: (2025)
di: Barin-Pacela, Vitoria, et al.
Pubblicazione: (2025)
Matryoshka Quantization
di: Nair, Pranav, et al.
Pubblicazione: (2025)
di: Nair, Pranav, et al.
Pubblicazione: (2025)
How to Parameterize Asymmetric Quantization Ranges for Quantization-Aware Training
di: You, Jaeseong, et al.
Pubblicazione: (2024)
di: You, Jaeseong, et al.
Pubblicazione: (2024)
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
di: Hooper, Coleman, et al.
Pubblicazione: (2024)
di: Hooper, Coleman, et al.
Pubblicazione: (2024)
Efficient Evaluation of Quantization-Effects in Neural Codecs
di: Mack, Wolfgang, et al.
Pubblicazione: (2025)
di: Mack, Wolfgang, et al.
Pubblicazione: (2025)
LPCD: Unified Framework from Layer-Wise to Submodule Quantization
di: Ichikawa, Yuma, et al.
Pubblicazione: (2025)
di: Ichikawa, Yuma, et al.
Pubblicazione: (2025)
Online Vector Quantized Attention
di: Alonso, Nick, et al.
Pubblicazione: (2026)
di: Alonso, Nick, et al.
Pubblicazione: (2026)
On the Spectral Flattening of Quantized Embeddings
di: Huang, Junlin, et al.
Pubblicazione: (2026)
di: Huang, Junlin, et al.
Pubblicazione: (2026)
Pyramid Vector Quantization for LLMs
di: van der Ouderaa, Tycho F. A., et al.
Pubblicazione: (2024)
di: van der Ouderaa, Tycho F. A., et al.
Pubblicazione: (2024)
CodeQuant: Unified Clustering and Quantization for Enhanced Outlier Smoothing in Low-Precision Mixture-of-Experts
di: Yin, Xiangyang, et al.
Pubblicazione: (2026)
di: Yin, Xiangyang, et al.
Pubblicazione: (2026)
UniQL: Unified Quantization and Low-rank Compression for Adaptive Edge LLMs
di: Chiang, Hung-Yueh, et al.
Pubblicazione: (2025)
di: Chiang, Hung-Yueh, et al.
Pubblicazione: (2025)
Preserve-Then-Quantize: Balancing Rank Budgets for Quantization Error Reconstruction in LLMs
di: Cho, Yoonjun, et al.
Pubblicazione: (2026)
di: Cho, Yoonjun, et al.
Pubblicazione: (2026)
Quantization Error Propagation: Revisiting Layer-Wise Post-Training Quantization
di: Arai, Yamato, et al.
Pubblicazione: (2025)
di: Arai, Yamato, et al.
Pubblicazione: (2025)
Divide-or-Conquer? Which Part Should You Distill Your LLM?
di: Wu, Zhuofeng, et al.
Pubblicazione: (2024)
di: Wu, Zhuofeng, et al.
Pubblicazione: (2024)
Efficient Post-training Quantization with FP8 Formats
di: Shen, Haihao, et al.
Pubblicazione: (2023)
di: Shen, Haihao, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Forensic Implications of Localized AI: Artifact Analysis of Ollama, LM Studio, and llama.cpp
di: Murtuza, Shariq
Pubblicazione: (2026) -
Pruning vs Quantization: Which is Better?
di: Kuzmin, Andrey, et al.
Pubblicazione: (2023) -
Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders
di: He, Zhengfu, et al.
Pubblicazione: (2024) -
Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct
di: Ackerman, Christopher, et al.
Pubblicazione: (2024) -
Domain Adaptation of Llama3-70B-Instruct through Continual Pre-Training and Model Merging: A Comprehensive Evaluation
di: Siriwardhana, Shamane, et al.
Pubblicazione: (2024)