From Bits to Chips: An LLM-based Hardware-Aware Quantization Agent for Streamlined Deployment of LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Deng, Kaiyuan, Zheng, Hangyu, Qing, Minghai, Zhu, Kunxiong, Li, Gen, Xiao, Yang, Zhang, Lan Emily, Guo, Linke, Hui, Bo, Wang, Yanzhi, Yuan, Geng, Agrawal, Gagan, Niu, Wei, Ma, Xiaolong
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915914535927808
author Deng, Kaiyuan
Zheng, Hangyu
Qing, Minghai
Zhu, Kunxiong
Li, Gen
Xiao, Yang
Zhang, Lan Emily
Guo, Linke
Hui, Bo
Wang, Yanzhi
Yuan, Geng
Agrawal, Gagan
Niu, Wei
Ma, Xiaolong
author_facet Deng, Kaiyuan
Zheng, Hangyu
Qing, Minghai
Zhu, Kunxiong
Li, Gen
Xiao, Yang
Zhang, Lan Emily
Guo, Linke
Hui, Bo
Wang, Yanzhi
Yuan, Geng
Agrawal, Gagan
Niu, Wei
Ma, Xiaolong
contents Deploying models, especially large language models (LLMs), is becoming increasingly attractive to a broader user base, including those without specialized expertise. However, due to the resource constraints of certain hardware, maintaining high accuracy with larger model while meeting the hardware requirements remains a significant challenge. Model quantization technique helps mitigate memory and compute bottlenecks, yet the added complexities of tuning and deploying quantized models further exacerbates these challenges, making the process unfriendly to most of the users. We introduce the Hardware-Aware Quantization Agent (HAQA), an automated framework that leverages LLMs to streamline the entire quantization and deployment process by enabling efficient hyperparameter tuning and hardware configuration, thereby simultaneously improving deployment quality and ease of use for a broad range of users. Our results demonstrate up to a 2.3x speedup in inference, along with increased throughput and improved accuracy compared to unoptimized models on Llama. Additionally, HAQA is designed to implement adaptive quantization strategies across diverse hardware platforms, as it automatically finds optimal settings even when they appear counterintuitive, thereby reducing extensive manual effort and demonstrating superior adaptability. Code will be released.
format Preprint
id arxiv_https___arxiv_org_abs_2601_03484
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle From Bits to Chips: An LLM-based Hardware-Aware Quantization Agent for Streamlined Deployment of LLMs
Deng, Kaiyuan
Zheng, Hangyu
Qing, Minghai
Zhu, Kunxiong
Li, Gen
Xiao, Yang
Zhang, Lan Emily
Guo, Linke
Hui, Bo
Wang, Yanzhi
Yuan, Geng
Agrawal, Gagan
Niu, Wei
Ma, Xiaolong
Machine Learning
Deploying models, especially large language models (LLMs), is becoming increasingly attractive to a broader user base, including those without specialized expertise. However, due to the resource constraints of certain hardware, maintaining high accuracy with larger model while meeting the hardware requirements remains a significant challenge. Model quantization technique helps mitigate memory and compute bottlenecks, yet the added complexities of tuning and deploying quantized models further exacerbates these challenges, making the process unfriendly to most of the users. We introduce the Hardware-Aware Quantization Agent (HAQA), an automated framework that leverages LLMs to streamline the entire quantization and deployment process by enabling efficient hyperparameter tuning and hardware configuration, thereby simultaneously improving deployment quality and ease of use for a broad range of users. Our results demonstrate up to a 2.3x speedup in inference, along with increased throughput and improved accuracy compared to unoptimized models on Llama. Additionally, HAQA is designed to implement adaptive quantization strategies across diverse hardware platforms, as it automatically finds optimal settings even when they appear counterintuitive, thereby reducing extensive manual effort and demonstrating superior adaptability. Code will be released.
title From Bits to Chips: An LLM-based Hardware-Aware Quantization Agent for Streamlined Deployment of LLMs
topic Machine Learning
url https://arxiv.org/abs/2601.03484