Hardware Acceleration of Kolmogorov-Arnold Network (KAN) for Lightweight Edge Inference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Wei-Hsing, Jia, Jianwei, Kong, Yuyao, Waqar, Faaiq, Wen, Tai-Hao, Chang, Meng-Fan, Yu, Shimeng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910609411407872
author Huang, Wei-Hsing
Jia, Jianwei
Kong, Yuyao
Waqar, Faaiq
Wen, Tai-Hao
Chang, Meng-Fan
Yu, Shimeng
author_facet Huang, Wei-Hsing
Jia, Jianwei
Kong, Yuyao
Waqar, Faaiq
Wen, Tai-Hao
Chang, Meng-Fan
Yu, Shimeng
contents Recently, a novel model named Kolmogorov-Arnold Networks (KAN) has been proposed with the potential to achieve the functionality of traditional deep neural networks (DNNs) using orders of magnitude fewer parameters by parameterized B-spline functions with trainable coefficients. However, the B-spline functions in KAN present new challenges for hardware acceleration. Evaluating the B-spline functions can be performed by using look-up tables (LUTs) to directly map the B-spline functions, thereby reducing computational resource requirements. However, this method still requires substantial circuit resources (LUTs, MUXs, decoders, etc.). For the first time, this paper employs an algorithm-hardware co-design methodology to accelerate KAN. The proposed algorithm-level techniques include Alignment-Symmetry and PowerGap KAN hardware aware quantization, KAN sparsity aware mapping strategy, and circuit-level techniques include N:1 Time Modulation Dynamic Voltage input generator with analog-CIM (ACIM) circuits. The impact of non-ideal effects, such as partial sum errors caused by the process variations, has been evaluated with the statistics measured from the TSMC 22nm RRAM-ACIM prototype chips. With the best searched hyperparameters of KAN and the optimized circuits implemented in 22 nm node, we can reduce hardware area by 41.78x, energy by 77.97x with 3.03% accuracy boost compared to the traditional DNN hardware.
format Preprint
id arxiv_https___arxiv_org_abs_2409_11418
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Hardware Acceleration of Kolmogorov-Arnold Network (KAN) for Lightweight Edge Inference
Huang, Wei-Hsing
Jia, Jianwei
Kong, Yuyao
Waqar, Faaiq
Wen, Tai-Hao
Chang, Meng-Fan
Yu, Shimeng
Hardware Architecture
Recently, a novel model named Kolmogorov-Arnold Networks (KAN) has been proposed with the potential to achieve the functionality of traditional deep neural networks (DNNs) using orders of magnitude fewer parameters by parameterized B-spline functions with trainable coefficients. However, the B-spline functions in KAN present new challenges for hardware acceleration. Evaluating the B-spline functions can be performed by using look-up tables (LUTs) to directly map the B-spline functions, thereby reducing computational resource requirements. However, this method still requires substantial circuit resources (LUTs, MUXs, decoders, etc.). For the first time, this paper employs an algorithm-hardware co-design methodology to accelerate KAN. The proposed algorithm-level techniques include Alignment-Symmetry and PowerGap KAN hardware aware quantization, KAN sparsity aware mapping strategy, and circuit-level techniques include N:1 Time Modulation Dynamic Voltage input generator with analog-CIM (ACIM) circuits. The impact of non-ideal effects, such as partial sum errors caused by the process variations, has been evaluated with the statistics measured from the TSMC 22nm RRAM-ACIM prototype chips. With the best searched hyperparameters of KAN and the optimized circuits implemented in 22 nm node, we can reduce hardware area by 41.78x, energy by 77.97x with 3.03% accuracy boost compared to the traditional DNN hardware.
title Hardware Acceleration of Kolmogorov-Arnold Network (KAN) for Lightweight Edge Inference
topic Hardware Architecture
url https://arxiv.org/abs/2409.11418