Mapping Space Exploration for Multi-Chiplet Accelerators Targeting LLM Inference Serving Workloads

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Boyu, Zhu, Zongwei, Xiong, Yi, Cao, Qianyue, Geng, Jiawei, Zhang, Xiaonan, Li, Xi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917376585367552
author Li, Boyu
Zhu, Zongwei
Xiong, Yi
Cao, Qianyue
Geng, Jiawei
Zhang, Xiaonan
Li, Xi
author_facet Li, Boyu
Zhu, Zongwei
Xiong, Yi
Cao, Qianyue
Geng, Jiawei
Zhang, Xiaonan
Li, Xi
contents Large Language Models (LLMs) impose massive computational demands, driving the need for scalable multi-chiplet accelerators. However, existing mapping space exploration efforts for such accelerators primarily focus on traditional CNN/Transformer workloads and fail to adequately support the dynamic behaviors of mixed request types and variable sequence lengths in real-world LLM inference serving. To bridge this gap, we first propose a computation execution graph-based mapping encoding scheme that decouples micro-batches and layers, enabling fine-grained execution control on heterogeneous chiplets and flexibly representing various parallelism strategies. Second, building upon this scheme, we develop the Compass framework, which integrates an evaluation engine and a genetic algorithm-based mapping generation engine to achieve efficient mapping search. Compared to state-of-the-art works, our solution achieves an average EDP reduction of 63.12%.
format Preprint
id arxiv_https___arxiv_org_abs_2512_06093
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mapping Space Exploration for Multi-Chiplet Accelerators Targeting LLM Inference Serving Workloads
Li, Boyu
Zhu, Zongwei
Xiong, Yi
Cao, Qianyue
Geng, Jiawei
Zhang, Xiaonan
Li, Xi
Hardware Architecture
Large Language Models (LLMs) impose massive computational demands, driving the need for scalable multi-chiplet accelerators. However, existing mapping space exploration efforts for such accelerators primarily focus on traditional CNN/Transformer workloads and fail to adequately support the dynamic behaviors of mixed request types and variable sequence lengths in real-world LLM inference serving. To bridge this gap, we first propose a computation execution graph-based mapping encoding scheme that decouples micro-batches and layers, enabling fine-grained execution control on heterogeneous chiplets and flexibly representing various parallelism strategies. Second, building upon this scheme, we develop the Compass framework, which integrates an evaluation engine and a genetic algorithm-based mapping generation engine to achieve efficient mapping search. Compared to state-of-the-art works, our solution achieves an average EDP reduction of 63.12%.
title Mapping Space Exploration for Multi-Chiplet Accelerators Targeting LLM Inference Serving Workloads
topic Hardware Architecture
url https://arxiv.org/abs/2512.06093