FusionCIM: Accelerating LLM Inference with Fusion-Driven Computing-in-Memory Architecture

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xuan, Zihao, Chen, Jia, Li, Yewen, Xuan, Wei, Chen, Hegan, Huo, Xiao, Tu, Fengbin
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917442971762688
author Xuan, Zihao
Chen, Jia
Li, Yewen
Xuan, Wei
Chen, Hegan
Huo, Xiao
Tu, Fengbin
author_facet Xuan, Zihao
Chen, Jia
Li, Yewen
Xuan, Wei
Chen, Hegan
Huo, Xiao
Tu, Fengbin
contents In this paper, we propose FusionCIM, an operator-fusion-driven compute-in-memory (CIM) accelerator architecture for efficient and scalable LLM inference, with three key innovations: (1) a hybrid CIM pipeline architecture that maps QKT computation on inner-product-based CIM (IP-CIM) and PV aggregation on outer-product-based CIM (OP-CIM) for efficient matrix multiplications fusion; (2) a QO-stationary dataflow that eliminates repeated KV loading in CIM and K-matrix access in buffer under transpose fusion, significantly improving data reuse on chip; and (3) a pattern-aware online-softmax mechanism that exploits distribution regularities of attention scores to reduce exponential rescaling overhead for non-linear fusion. Experimental results on LLaMA-3 model show that FusionCIM achieves up to 3.86x energy saving, and 1.98x speedup compared with prior SOTA CIM-based designs with 29.4 TOPS/W energy efficiency at the system level.
format Preprint
id arxiv_https___arxiv_org_abs_2604_25317
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FusionCIM: Accelerating LLM Inference with Fusion-Driven Computing-in-Memory Architecture
Xuan, Zihao
Chen, Jia
Li, Yewen
Xuan, Wei
Chen, Hegan
Huo, Xiao
Tu, Fengbin
Hardware Architecture
In this paper, we propose FusionCIM, an operator-fusion-driven compute-in-memory (CIM) accelerator architecture for efficient and scalable LLM inference, with three key innovations: (1) a hybrid CIM pipeline architecture that maps QKT computation on inner-product-based CIM (IP-CIM) and PV aggregation on outer-product-based CIM (OP-CIM) for efficient matrix multiplications fusion; (2) a QO-stationary dataflow that eliminates repeated KV loading in CIM and K-matrix access in buffer under transpose fusion, significantly improving data reuse on chip; and (3) a pattern-aware online-softmax mechanism that exploits distribution regularities of attention scores to reduce exponential rescaling overhead for non-linear fusion. Experimental results on LLaMA-3 model show that FusionCIM achieves up to 3.86x energy saving, and 1.98x speedup compared with prior SOTA CIM-based designs with 29.4 TOPS/W energy efficiency at the system level.
title FusionCIM: Accelerating LLM Inference with Fusion-Driven Computing-in-Memory Architecture
topic Hardware Architecture
url https://arxiv.org/abs/2604.25317