FusionCIM: Accelerating LLM Inference with Fusion-Driven Computing-in-Memory Architecture
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866917442971762688 |
|---|---|
| author | Xuan, Zihao Chen, Jia Li, Yewen Xuan, Wei Chen, Hegan Huo, Xiao Tu, Fengbin |
| author_facet | Xuan, Zihao Chen, Jia Li, Yewen Xuan, Wei Chen, Hegan Huo, Xiao Tu, Fengbin |
| contents | In this paper, we propose FusionCIM, an operator-fusion-driven compute-in-memory (CIM) accelerator architecture for efficient and scalable LLM inference, with three key innovations: (1) a hybrid CIM pipeline architecture that maps QKT computation on inner-product-based CIM (IP-CIM) and PV aggregation on outer-product-based CIM (OP-CIM) for efficient matrix multiplications fusion; (2) a QO-stationary dataflow that eliminates repeated KV loading in CIM and K-matrix access in buffer under transpose fusion, significantly improving data reuse on chip; and (3) a pattern-aware online-softmax mechanism that exploits distribution regularities of attention scores to reduce exponential rescaling overhead for non-linear fusion. Experimental results on LLaMA-3 model show that FusionCIM achieves up to 3.86x energy saving, and 1.98x speedup compared with prior SOTA CIM-based designs with 29.4 TOPS/W energy efficiency at the system level. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_25317 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | FusionCIM: Accelerating LLM Inference with Fusion-Driven Computing-in-Memory Architecture Xuan, Zihao Chen, Jia Li, Yewen Xuan, Wei Chen, Hegan Huo, Xiao Tu, Fengbin Hardware Architecture In this paper, we propose FusionCIM, an operator-fusion-driven compute-in-memory (CIM) accelerator architecture for efficient and scalable LLM inference, with three key innovations: (1) a hybrid CIM pipeline architecture that maps QKT computation on inner-product-based CIM (IP-CIM) and PV aggregation on outer-product-based CIM (OP-CIM) for efficient matrix multiplications fusion; (2) a QO-stationary dataflow that eliminates repeated KV loading in CIM and K-matrix access in buffer under transpose fusion, significantly improving data reuse on chip; and (3) a pattern-aware online-softmax mechanism that exploits distribution regularities of attention scores to reduce exponential rescaling overhead for non-linear fusion. Experimental results on LLaMA-3 model show that FusionCIM achieves up to 3.86x energy saving, and 1.98x speedup compared with prior SOTA CIM-based designs with 29.4 TOPS/W energy efficiency at the system level. |
| title | FusionCIM: Accelerating LLM Inference with Fusion-Driven Computing-in-Memory Architecture |
| topic | Hardware Architecture |
| url | https://arxiv.org/abs/2604.25317 |