ONNXim: A Fast, Cycle-level Multi-core NPU Simulator

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ham, Hyungkyu, Yang, Wonhyuk, Shin, Yunseon, Woo, Okkyun, Heo, Guseul, Lee, Sangyeop, Park, Jongse, Kim, Gwangsun
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910483001376768
author Ham, Hyungkyu
Yang, Wonhyuk
Shin, Yunseon
Woo, Okkyun
Heo, Guseul
Lee, Sangyeop
Park, Jongse
Kim, Gwangsun
author_facet Ham, Hyungkyu
Yang, Wonhyuk
Shin, Yunseon
Woo, Okkyun
Heo, Guseul
Lee, Sangyeop
Park, Jongse
Kim, Gwangsun
contents As DNNs are widely adopted in various application domains while demanding increasingly higher compute and memory requirements, designing efficient and performant NPUs (Neural Processing Units) is becoming more important. However, existing architectural NPU simulators lack support for high-speed simulation, multi-core modeling, multi-tenant scenarios, detailed DRAM/NoC modeling, and/or different deep learning frameworks. To address these limitations, this work proposes ONNXim, a fast cycle-level simulator for multi-core NPUs in DNN serving systems. It takes DNN models represented in the ONNX graph format generated from various deep learning frameworks for ease of simulation. In addition, based on the observation that typical NPU cores process tensor tiles from on-chip scratchpad memory with deterministic compute latency, we forgo a detailed modeling for the computation while still preserving simulation accuracy. ONNXim also preserves dependencies between compute and tile DMAs. Meanwhile, the DRAM and NoC are modeled in cycle-level to properly model contention among multiple cores that can execute different DNN models for multi-tenancy. Consequently, ONNXim is significantly faster than existing simulators (e.g., by up to 384x over Accel-sim) and enables various case studies, such as multi-tenant NPUs, that were previously impractical due to slow speed and/or lack of functionalities. ONNXim is publicly available at https://github.com/PSAL-POSTECH/ONNXim.
format Preprint
id arxiv_https___arxiv_org_abs_2406_08051
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ONNXim: A Fast, Cycle-level Multi-core NPU Simulator
Ham, Hyungkyu
Yang, Wonhyuk
Shin, Yunseon
Woo, Okkyun
Heo, Guseul
Lee, Sangyeop
Park, Jongse
Kim, Gwangsun
Hardware Architecture
Performance
As DNNs are widely adopted in various application domains while demanding increasingly higher compute and memory requirements, designing efficient and performant NPUs (Neural Processing Units) is becoming more important. However, existing architectural NPU simulators lack support for high-speed simulation, multi-core modeling, multi-tenant scenarios, detailed DRAM/NoC modeling, and/or different deep learning frameworks. To address these limitations, this work proposes ONNXim, a fast cycle-level simulator for multi-core NPUs in DNN serving systems. It takes DNN models represented in the ONNX graph format generated from various deep learning frameworks for ease of simulation. In addition, based on the observation that typical NPU cores process tensor tiles from on-chip scratchpad memory with deterministic compute latency, we forgo a detailed modeling for the computation while still preserving simulation accuracy. ONNXim also preserves dependencies between compute and tile DMAs. Meanwhile, the DRAM and NoC are modeled in cycle-level to properly model contention among multiple cores that can execute different DNN models for multi-tenancy. Consequently, ONNXim is significantly faster than existing simulators (e.g., by up to 384x over Accel-sim) and enables various case studies, such as multi-tenant NPUs, that were previously impractical due to slow speed and/or lack of functionalities. ONNXim is publicly available at https://github.com/PSAL-POSTECH/ONNXim.
title ONNXim: A Fast, Cycle-level Multi-core NPU Simulator
topic Hardware Architecture
Performance
url https://arxiv.org/abs/2406.08051