Performance Implications of Multi-Chiplet Neural Processing Units on Autonomous Driving Perception
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866916494214955008 |
|---|---|
| author | Odema, Mohanad Chen, Luke Kwon, Hyoukjun Faruque, Mohammad Abdullah Al |
| author_facet | Odema, Mohanad Chen, Luke Kwon, Hyoukjun Faruque, Mohammad Abdullah Al |
| contents | We study the application of emerging chiplet-based Neural Processing Units to accelerate vehicular AI perception workloads in constrained automotive settings. The motivation stems from how chiplets technology is becoming integral to emerging vehicular architectures, providing a cost-effective trade-off between performance, modularity, and customization; and from perception models being the most computationally demanding workloads in a autonomous driving system. Using the Tesla Autopilot perception pipeline as a case study, we first breakdown its constituent models and profile their performance on different chiplet accelerators. From the insights, we propose a novel scheduling strategy to efficiently deploy perception workloads on multi-chip AI accelerators. Our experiments using a standard DNN performance simulator, MAESTRO, show our approach realizes 82% and 2.8x increase in throughput and processing engines utilization compared to monolithic accelerator designs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_16007 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Performance Implications of Multi-Chiplet Neural Processing Units on Autonomous Driving Perception Odema, Mohanad Chen, Luke Kwon, Hyoukjun Faruque, Mohammad Abdullah Al Hardware Architecture Artificial Intelligence Distributed, Parallel, and Cluster Computing Performance We study the application of emerging chiplet-based Neural Processing Units to accelerate vehicular AI perception workloads in constrained automotive settings. The motivation stems from how chiplets technology is becoming integral to emerging vehicular architectures, providing a cost-effective trade-off between performance, modularity, and customization; and from perception models being the most computationally demanding workloads in a autonomous driving system. Using the Tesla Autopilot perception pipeline as a case study, we first breakdown its constituent models and profile their performance on different chiplet accelerators. From the insights, we propose a novel scheduling strategy to efficiently deploy perception workloads on multi-chip AI accelerators. Our experiments using a standard DNN performance simulator, MAESTRO, show our approach realizes 82% and 2.8x increase in throughput and processing engines utilization compared to monolithic accelerator designs. |
| title | Performance Implications of Multi-Chiplet Neural Processing Units on Autonomous Driving Perception |
| topic | Hardware Architecture Artificial Intelligence Distributed, Parallel, and Cluster Computing Performance |
| url | https://arxiv.org/abs/2411.16007 |