Cambricon-LLM: A Chiplet-Based Hybrid Architecture for On-Device Inference of 70B LLM

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Zhongkai, Liang, Shengwen, Ma, Tianyun, Cai, Yunke, Nan, Ziyuan, Huang, Di, Song, Xinkai, Hao, Yifan, Zhang, Jie, Zhi, Tian, Zhao, Yongwei, Du, Zidong, Hu, Xing, Guo, Qi, Chen, Tianshi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910617983516672
author Yu, Zhongkai
Liang, Shengwen
Ma, Tianyun
Cai, Yunke
Nan, Ziyuan
Huang, Di
Song, Xinkai
Hao, Yifan
Zhang, Jie
Zhi, Tian
Zhao, Yongwei
Du, Zidong
Hu, Xing
Guo, Qi
Chen, Tianshi
author_facet Yu, Zhongkai
Liang, Shengwen
Ma, Tianyun
Cai, Yunke
Nan, Ziyuan
Huang, Di
Song, Xinkai
Hao, Yifan
Zhang, Jie
Zhi, Tian
Zhao, Yongwei
Du, Zidong
Hu, Xing
Guo, Qi
Chen, Tianshi
contents Deploying advanced large language models on edge devices, such as smartphones and robotics, is a growing trend that enhances user data privacy and network connectivity resilience while preserving intelligent capabilities. However, such a task exhibits single-batch computing with incredibly low arithmetic intensity, which poses the significant challenges of huge memory footprint and bandwidth demands on limited edge resources. To address these issues, we introduce Cambricon-LLM, a chiplet-based hybrid architecture with NPU and a dedicated NAND flash chip to enable efficient on-device inference of 70B LLMs. Such a hybrid architecture utilizes both the high computing capability of NPU and the data capacity of the NAND flash chip, with the proposed hardware-tiling strategy that minimizes the data movement overhead between NPU and NAND flash chip. Specifically, the NAND flash chip, enhanced by our innovative in-flash computing and on-die ECC techniques, excels at performing precise lightweight on-die processing. Simultaneously, the NPU collaborates with the flash chip for matrix operations and handles special function computations beyond the flash's on-die processing capabilities. Overall, Cambricon-LLM enables the on-device inference of 70B LLMs at a speed of 3.44 token/s, and 7B LLMs at a speed of 36.34 token/s, which is over 22X to 45X faster than existing flash-offloading technologies, showing the potentiality of deploying powerful LLMs in edge devices.
format Preprint
id arxiv_https___arxiv_org_abs_2409_15654
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Cambricon-LLM: A Chiplet-Based Hybrid Architecture for On-Device Inference of 70B LLM
Yu, Zhongkai
Liang, Shengwen
Ma, Tianyun
Cai, Yunke
Nan, Ziyuan
Huang, Di
Song, Xinkai
Hao, Yifan
Zhang, Jie
Zhi, Tian
Zhao, Yongwei
Du, Zidong
Hu, Xing
Guo, Qi
Chen, Tianshi
Hardware Architecture
Deploying advanced large language models on edge devices, such as smartphones and robotics, is a growing trend that enhances user data privacy and network connectivity resilience while preserving intelligent capabilities. However, such a task exhibits single-batch computing with incredibly low arithmetic intensity, which poses the significant challenges of huge memory footprint and bandwidth demands on limited edge resources. To address these issues, we introduce Cambricon-LLM, a chiplet-based hybrid architecture with NPU and a dedicated NAND flash chip to enable efficient on-device inference of 70B LLMs. Such a hybrid architecture utilizes both the high computing capability of NPU and the data capacity of the NAND flash chip, with the proposed hardware-tiling strategy that minimizes the data movement overhead between NPU and NAND flash chip. Specifically, the NAND flash chip, enhanced by our innovative in-flash computing and on-die ECC techniques, excels at performing precise lightweight on-die processing. Simultaneously, the NPU collaborates with the flash chip for matrix operations and handles special function computations beyond the flash's on-die processing capabilities. Overall, Cambricon-LLM enables the on-device inference of 70B LLMs at a speed of 3.44 token/s, and 7B LLMs at a speed of 36.34 token/s, which is over 22X to 45X faster than existing flash-offloading technologies, showing the potentiality of deploying powerful LLMs in edge devices.
title Cambricon-LLM: A Chiplet-Based Hybrid Architecture for On-Device Inference of 70B LLM
topic Hardware Architecture
url https://arxiv.org/abs/2409.15654