Performance Implications of Multi-Chiplet Neural Processing Units on Autonomous Driving Perception

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Odema, Mohanad, Chen, Luke, Kwon, Hyoukjun, Faruque, Mohammad Abdullah Al
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916494214955008
author Odema, Mohanad
Chen, Luke
Kwon, Hyoukjun
Faruque, Mohammad Abdullah Al
author_facet Odema, Mohanad
Chen, Luke
Kwon, Hyoukjun
Faruque, Mohammad Abdullah Al
contents We study the application of emerging chiplet-based Neural Processing Units to accelerate vehicular AI perception workloads in constrained automotive settings. The motivation stems from how chiplets technology is becoming integral to emerging vehicular architectures, providing a cost-effective trade-off between performance, modularity, and customization; and from perception models being the most computationally demanding workloads in a autonomous driving system. Using the Tesla Autopilot perception pipeline as a case study, we first breakdown its constituent models and profile their performance on different chiplet accelerators. From the insights, we propose a novel scheduling strategy to efficiently deploy perception workloads on multi-chip AI accelerators. Our experiments using a standard DNN performance simulator, MAESTRO, show our approach realizes 82% and 2.8x increase in throughput and processing engines utilization compared to monolithic accelerator designs.
format Preprint
id arxiv_https___arxiv_org_abs_2411_16007
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Performance Implications of Multi-Chiplet Neural Processing Units on Autonomous Driving Perception
Odema, Mohanad
Chen, Luke
Kwon, Hyoukjun
Faruque, Mohammad Abdullah Al
Hardware Architecture
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Performance
We study the application of emerging chiplet-based Neural Processing Units to accelerate vehicular AI perception workloads in constrained automotive settings. The motivation stems from how chiplets technology is becoming integral to emerging vehicular architectures, providing a cost-effective trade-off between performance, modularity, and customization; and from perception models being the most computationally demanding workloads in a autonomous driving system. Using the Tesla Autopilot perception pipeline as a case study, we first breakdown its constituent models and profile their performance on different chiplet accelerators. From the insights, we propose a novel scheduling strategy to efficiently deploy perception workloads on multi-chip AI accelerators. Our experiments using a standard DNN performance simulator, MAESTRO, show our approach realizes 82% and 2.8x increase in throughput and processing engines utilization compared to monolithic accelerator designs.
title Performance Implications of Multi-Chiplet Neural Processing Units on Autonomous Driving Perception
topic Hardware Architecture
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Performance
url https://arxiv.org/abs/2411.16007