A Tensor-Train Decomposition based Compression of LLMs on Group Vector Systolic Accelerator

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Sixiao, Wang, Tintin, Li, Ang, Shen, Ao, Li, Kai, Jiang, Keyao, Huang, Mingqiang, Yu, Hao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917907637731328
author Huang, Sixiao
Wang, Tintin
Li, Ang
Shen, Ao
Li, Kai
Jiang, Keyao
Huang, Mingqiang
Yu, Hao
author_facet Huang, Sixiao
Wang, Tintin
Li, Ang
Shen, Ao
Li, Kai
Jiang, Keyao
Huang, Mingqiang
Yu, Hao
contents Large language models (LLMs) are both storage-intensive and computation-intensive, posing significant challenges when deployed on resource-constrained hardware. As linear layers in LLMs are mainly resource consuming parts, this paper develops a tensor-train decomposition (TTD) for LLMs with a further hardware implementation on FPGA. TTD compression is applied to the linear layers in ChatGLM3-6B and LLaMA2-7B models with compression ratios (CRs) for the whole network 1.94$\times$ and 1.60$\times$, respectively. The compressed LLMs are further implemented on FPGA hardware within a highly efficient group vector systolic array (GVSA) architecture, which has DSP-shared parallel vector PEs for TTD inference, as well as optimized data communication in the accelerator. Experimental results show that the corresponding TTD based LLM accelerator implemented on FPGA achieves 1.45$\times$ and 1.57$\times$ reduction in first token delay for ChatGLM3-6B and LLaMA2-7B models, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2501_19135
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Tensor-Train Decomposition based Compression of LLMs on Group Vector Systolic Accelerator
Huang, Sixiao
Wang, Tintin
Li, Ang
Shen, Ao
Li, Kai
Jiang, Keyao
Huang, Mingqiang
Yu, Hao
Hardware Architecture
Large language models (LLMs) are both storage-intensive and computation-intensive, posing significant challenges when deployed on resource-constrained hardware. As linear layers in LLMs are mainly resource consuming parts, this paper develops a tensor-train decomposition (TTD) for LLMs with a further hardware implementation on FPGA. TTD compression is applied to the linear layers in ChatGLM3-6B and LLaMA2-7B models with compression ratios (CRs) for the whole network 1.94$\times$ and 1.60$\times$, respectively. The compressed LLMs are further implemented on FPGA hardware within a highly efficient group vector systolic array (GVSA) architecture, which has DSP-shared parallel vector PEs for TTD inference, as well as optimized data communication in the accelerator. Experimental results show that the corresponding TTD based LLM accelerator implemented on FPGA achieves 1.45$\times$ and 1.57$\times$ reduction in first token delay for ChatGLM3-6B and LLaMA2-7B models, respectively.
title A Tensor-Train Decomposition based Compression of LLMs on Group Vector Systolic Accelerator
topic Hardware Architecture
url https://arxiv.org/abs/2501.19135