A dynamic parallel method for performance optimization on hybrid CPUs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yu, Luo, Yucheng, Liu, Haihao, Shen
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929608847261696
author Yu, Luo
Yucheng, Liu
Haihao, Shen
author_facet Yu, Luo
Yucheng, Liu
Haihao, Shen
contents The AIPC concept is gaining popularity, and more and more hybrid CPUs will be running AI models on client devices. However, the current AI inference framework overlooks the imbalanced hardware capability of hybrid CPUs, leading to low inference performance. To address this issue, we have introduced a dynamic parallel method for hybrid CPUs, which significantly increases LLM inference performance by balancing the workload for each core of a hybrid CPU before the parallel work starts. This method has enabled Neural Speed to achieve more than 90% (on average) of memory bandwidth on two hybrid Intel CPUs.
format Preprint
id arxiv_https___arxiv_org_abs_2411_19542
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A dynamic parallel method for performance optimization on hybrid CPUs
Yu, Luo
Yucheng, Liu
Haihao, Shen
Distributed, Parallel, and Cluster Computing
Performance
The AIPC concept is gaining popularity, and more and more hybrid CPUs will be running AI models on client devices. However, the current AI inference framework overlooks the imbalanced hardware capability of hybrid CPUs, leading to low inference performance. To address this issue, we have introduced a dynamic parallel method for hybrid CPUs, which significantly increases LLM inference performance by balancing the workload for each core of a hybrid CPU before the parallel work starts. This method has enabled Neural Speed to achieve more than 90% (on average) of memory bandwidth on two hybrid Intel CPUs.
title A dynamic parallel method for performance optimization on hybrid CPUs
topic Distributed, Parallel, and Cluster Computing
Performance
url https://arxiv.org/abs/2411.19542