Mapping Space Exploration for Multi-Chiplet Accelerators Targeting LLM Inference Serving Workloads
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917376585367552 |
|---|---|
| author | Li, Boyu Zhu, Zongwei Xiong, Yi Cao, Qianyue Geng, Jiawei Zhang, Xiaonan Li, Xi |
| author_facet | Li, Boyu Zhu, Zongwei Xiong, Yi Cao, Qianyue Geng, Jiawei Zhang, Xiaonan Li, Xi |
| contents | Large Language Models (LLMs) impose massive computational demands, driving the need for scalable multi-chiplet accelerators. However, existing mapping space exploration efforts for such accelerators primarily focus on traditional CNN/Transformer workloads and fail to adequately support the dynamic behaviors of mixed request types and variable sequence lengths in real-world LLM inference serving. To bridge this gap, we first propose a computation execution graph-based mapping encoding scheme that decouples micro-batches and layers, enabling fine-grained execution control on heterogeneous chiplets and flexibly representing various parallelism strategies. Second, building upon this scheme, we develop the Compass framework, which integrates an evaluation engine and a genetic algorithm-based mapping generation engine to achieve efficient mapping search. Compared to state-of-the-art works, our solution achieves an average EDP reduction of 63.12%. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_06093 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Mapping Space Exploration for Multi-Chiplet Accelerators Targeting LLM Inference Serving Workloads Li, Boyu Zhu, Zongwei Xiong, Yi Cao, Qianyue Geng, Jiawei Zhang, Xiaonan Li, Xi Hardware Architecture Large Language Models (LLMs) impose massive computational demands, driving the need for scalable multi-chiplet accelerators. However, existing mapping space exploration efforts for such accelerators primarily focus on traditional CNN/Transformer workloads and fail to adequately support the dynamic behaviors of mixed request types and variable sequence lengths in real-world LLM inference serving. To bridge this gap, we first propose a computation execution graph-based mapping encoding scheme that decouples micro-batches and layers, enabling fine-grained execution control on heterogeneous chiplets and flexibly representing various parallelism strategies. Second, building upon this scheme, we develop the Compass framework, which integrates an evaluation engine and a genetic algorithm-based mapping generation engine to achieve efficient mapping search. Compared to state-of-the-art works, our solution achieves an average EDP reduction of 63.12%. |
| title | Mapping Space Exploration for Multi-Chiplet Accelerators Targeting LLM Inference Serving Workloads |
| topic | Hardware Architecture |
| url | https://arxiv.org/abs/2512.06093 |