FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Park, Gunho, Kwon, Hyeokjun, Kim, Jiwoo, Bae, Jeongin, Park, Baeseong, Lee, Dongsoo, Lee, Youngjoo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913727486361600
author Park, Gunho
Kwon, Hyeokjun
Kim, Jiwoo
Bae, Jeongin
Park, Baeseong
Lee, Dongsoo
Lee, Youngjoo
author_facet Park, Gunho
Kwon, Hyeokjun
Kim, Jiwoo
Bae, Jeongin
Park, Baeseong
Lee, Dongsoo
Lee, Youngjoo
contents Weight-only quantization has emerged as a promising solution to the deployment challenges of large language models (LLMs). However, it necessitates FP-INT operations, which make implementation on general-purpose hardware like GPUs difficult. In this paper, we propose FIGLUT, an efficient look-up table (LUT)-based GEMM accelerator architecture. Instead of performing traditional arithmetic operations, FIGLUT retrieves precomputed values from an LUT based on weight patterns, significantly reducing the computational complexity. We also introduce a novel LUT design that addresses the limitations of conventional memory architectures. To further improve LUT-based operations, we propose a half-size LUT combined with a dedicated decoding and multiplexing unit. FIGLUT efficiently supports different bit precisions and quantization methods using a single fixed hardware configuration. For the same 3-bit weight precision, FIGLUT demonstrates 59% higher TOPS/W and 20% lower perplexity than state-of-the-art accelerator designs. When targeting the same perplexity, FIGLUT achieves 98% higher TOPS/W by performing 2.4-bit operations.
format Preprint
id arxiv_https___arxiv_org_abs_2503_06862
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables
Park, Gunho
Kwon, Hyeokjun
Kim, Jiwoo
Bae, Jeongin
Park, Baeseong
Lee, Dongsoo
Lee, Youngjoo
Hardware Architecture
Weight-only quantization has emerged as a promising solution to the deployment challenges of large language models (LLMs). However, it necessitates FP-INT operations, which make implementation on general-purpose hardware like GPUs difficult. In this paper, we propose FIGLUT, an efficient look-up table (LUT)-based GEMM accelerator architecture. Instead of performing traditional arithmetic operations, FIGLUT retrieves precomputed values from an LUT based on weight patterns, significantly reducing the computational complexity. We also introduce a novel LUT design that addresses the limitations of conventional memory architectures. To further improve LUT-based operations, we propose a half-size LUT combined with a dedicated decoding and multiplexing unit. FIGLUT efficiently supports different bit precisions and quantization methods using a single fixed hardware configuration. For the same 3-bit weight precision, FIGLUT demonstrates 59% higher TOPS/W and 20% lower perplexity than state-of-the-art accelerator designs. When targeting the same perplexity, FIGLUT achieves 98% higher TOPS/W by performing 2.4-bit operations.
title FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables
topic Hardware Architecture
url https://arxiv.org/abs/2503.06862