SmartQuant: CXL-based AI Model Store in Support of Runtime Configurable Weight Quantization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xie, Rui, Haq, Asad Ul, Ma, Linsen, Sun, Krystal, Sen, Sanchari, Venkataramani, Swagath, Liu, Liu, Zhang, Tong
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916360549826560
author Xie, Rui
Haq, Asad Ul
Ma, Linsen
Sun, Krystal
Sen, Sanchari
Venkataramani, Swagath
Liu, Liu
Zhang, Tong
author_facet Xie, Rui
Haq, Asad Ul
Ma, Linsen
Sun, Krystal
Sen, Sanchari
Venkataramani, Swagath
Liu, Liu
Zhang, Tong
contents Recent studies have revealed that, during the inference on generative AI models such as transformer, the importance of different weights exhibits substantial context-dependent variations. This naturally manifests a promising potential of adaptively configuring weight quantization to improve the generative AI inference efficiency. Although configurable weight quantization can readily leverage the hardware support of variable-precision arithmetics in modern GPU and AI accelerators, little prior research has studied how one could exploit variable weight quantization to proportionally improve the AI model memory access speed and energy efficiency. Motivated by the rapidly maturing CXL ecosystem, this work develops a CXL-based design solution to fill this gap. The key is to allow CXL memory controllers play an active role in supporting and exploiting runtime configurable weight quantization. Using transformer as a representative generative AI model, we carried out experiments that well demonstrate the effectiveness of the proposed design solution.
format Preprint
id arxiv_https___arxiv_org_abs_2407_15866
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SmartQuant: CXL-based AI Model Store in Support of Runtime Configurable Weight Quantization
Xie, Rui
Haq, Asad Ul
Ma, Linsen
Sun, Krystal
Sen, Sanchari
Venkataramani, Swagath
Liu, Liu
Zhang, Tong
Machine Learning
Artificial Intelligence
Hardware Architecture
Recent studies have revealed that, during the inference on generative AI models such as transformer, the importance of different weights exhibits substantial context-dependent variations. This naturally manifests a promising potential of adaptively configuring weight quantization to improve the generative AI inference efficiency. Although configurable weight quantization can readily leverage the hardware support of variable-precision arithmetics in modern GPU and AI accelerators, little prior research has studied how one could exploit variable weight quantization to proportionally improve the AI model memory access speed and energy efficiency. Motivated by the rapidly maturing CXL ecosystem, this work develops a CXL-based design solution to fill this gap. The key is to allow CXL memory controllers play an active role in supporting and exploiting runtime configurable weight quantization. Using transformer as a representative generative AI model, we carried out experiments that well demonstrate the effectiveness of the proposed design solution.
title SmartQuant: CXL-based AI Model Store in Support of Runtime Configurable Weight Quantization
topic Machine Learning
Artificial Intelligence
Hardware Architecture
url https://arxiv.org/abs/2407.15866