QUARK: Quantization-Enabled Circuit Sharing for Transformer Acceleration by Exploiting Common Patterns in Nonlinear Operations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhao, Zhixiong, Li, Haomin, Liu, Fangxin, Lu, Yuncheng, Wang, Zongwu, Yang, Tao, Jiang, Li, Guan, Haibing
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915888564797440
author Zhao, Zhixiong
Li, Haomin
Liu, Fangxin
Lu, Yuncheng
Wang, Zongwu
Yang, Tao
Jiang, Li
Guan, Haibing
author_facet Zhao, Zhixiong
Li, Haomin
Liu, Fangxin
Lu, Yuncheng
Wang, Zongwu
Yang, Tao
Jiang, Li
Guan, Haibing
contents Transformer-based models have revolutionized computer vision (CV) and natural language processing (NLP) by achieving state-of-the-art performance across a range of benchmarks. However, nonlinear operations in models significantly contribute to inference latency, presenting unique challenges for efficient hardware acceleration. To this end, we propose QUARK, a quantization-enabled FPGA acceleration framework that leverages common patterns in nonlinear operations to enable efficient circuit sharing, thereby reducing hardware resource requirements. QUARK targets all nonlinear operations within Transformer-based models, achieving high-performance approximation through a novel circuit-sharing design tailored to accelerate these operations. Our evaluation demonstrates that QUARK significantly reduces the computational overhead of nonlinear operators in mainstream Transformer architectures, achieving up to a 1.96 times end-to-end speedup over GPU implementations. Moreover, QUARK lowers the hardware overhead of nonlinear modules by more than 50% compared to prior approaches, all while maintaining high model accuracy -- and even substantially boosting accuracy under ultra-low-bit quantization.
format Preprint
id arxiv_https___arxiv_org_abs_2511_06767
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle QUARK: Quantization-Enabled Circuit Sharing for Transformer Acceleration by Exploiting Common Patterns in Nonlinear Operations
Zhao, Zhixiong
Li, Haomin
Liu, Fangxin
Lu, Yuncheng
Wang, Zongwu
Yang, Tao
Jiang, Li
Guan, Haibing
Machine Learning
Artificial Intelligence
Transformer-based models have revolutionized computer vision (CV) and natural language processing (NLP) by achieving state-of-the-art performance across a range of benchmarks. However, nonlinear operations in models significantly contribute to inference latency, presenting unique challenges for efficient hardware acceleration. To this end, we propose QUARK, a quantization-enabled FPGA acceleration framework that leverages common patterns in nonlinear operations to enable efficient circuit sharing, thereby reducing hardware resource requirements. QUARK targets all nonlinear operations within Transformer-based models, achieving high-performance approximation through a novel circuit-sharing design tailored to accelerate these operations. Our evaluation demonstrates that QUARK significantly reduces the computational overhead of nonlinear operators in mainstream Transformer architectures, achieving up to a 1.96 times end-to-end speedup over GPU implementations. Moreover, QUARK lowers the hardware overhead of nonlinear modules by more than 50% compared to prior approaches, all while maintaining high model accuracy -- and even substantially boosting accuracy under ultra-low-bit quantization.
title QUARK: Quantization-Enabled Circuit Sharing for Transformer Acceleration by Exploiting Common Patterns in Nonlinear Operations
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2511.06767