Sandwich: Joint Configuration Search and Hot-Switching for Efficient CPU LLM Serving

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhao, Juntao, Li, Jiuru, Wu, Chuan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918446642495488
author Zhao, Juntao
Li, Jiuru
Wu, Chuan
author_facet Zhao, Juntao
Li, Jiuru
Wu, Chuan
contents CPUs are critical for LLM serving due to their availability, cost efficiency, and edge applicability. However, efficient CPU serving is hindered by conflicting prefill/decode resource demands under non-disaggregated deployment constraints--existing solutions fail to avoid cross-phase interference, ignore sub-NUMA hardware structures, and deliver suboptimal dynamic-shape kernel performance. We propose Sandwich, a full-stack CPU LLM serving system with three core innovations addressing these challenges: (1) seamless phase-wise plan switching to eliminate cross-phase interference; (2) TopoTree, a tree-based hardware abstraction for automated substructure-aware (e.g., LLC slices) partial core allocation; (3) fast-start-then-finetune dynamic-shape tensor program generation. Across five x86/ARM CPU platforms, Sandwich achieves an average 2.01x end-to-end speedup and up to 3.40x latency reduction over state-of-the-art systems. Its kernels match static compiler performance with three orders of magnitude lower tuning cost.
format Preprint
id arxiv_https___arxiv_org_abs_2507_18454
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Sandwich: Joint Configuration Search and Hot-Switching for Efficient CPU LLM Serving
Zhao, Juntao
Li, Jiuru
Wu, Chuan
Hardware Architecture
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Programming Languages
CPUs are critical for LLM serving due to their availability, cost efficiency, and edge applicability. However, efficient CPU serving is hindered by conflicting prefill/decode resource demands under non-disaggregated deployment constraints--existing solutions fail to avoid cross-phase interference, ignore sub-NUMA hardware structures, and deliver suboptimal dynamic-shape kernel performance. We propose Sandwich, a full-stack CPU LLM serving system with three core innovations addressing these challenges: (1) seamless phase-wise plan switching to eliminate cross-phase interference; (2) TopoTree, a tree-based hardware abstraction for automated substructure-aware (e.g., LLC slices) partial core allocation; (3) fast-start-then-finetune dynamic-shape tensor program generation. Across five x86/ARM CPU platforms, Sandwich achieves an average 2.01x end-to-end speedup and up to 3.40x latency reduction over state-of-the-art systems. Its kernels match static compiler performance with three orders of magnitude lower tuning cost.
title Sandwich: Joint Configuration Search and Hot-Switching for Efficient CPU LLM Serving
topic Hardware Architecture
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Programming Languages
url https://arxiv.org/abs/2507.18454