Accelerating stencils on the Tenstorrent Grayskull RISC-V accelerator

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Brown, Nick, Barton, Ryan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912048469770240
author Brown, Nick
Barton, Ryan
author_facet Brown, Nick
Barton, Ryan
contents The RISC-V Instruction Set Architecture (ISA) has enjoyed phenomenal growth in recent years, however it still to gain popularity in HPC. Whilst adopting RISC-V CPU solutions in HPC might be some way off, RISC-V based PCIe accelerators offer a middle ground where vendors benefit from the flexibility of RISC-V yet fit into existing systems. In this paper we focus on the Tenstorrent Grayskull PCIe RISC-V based accelerator which, built upon Tensix cores, decouples data movement from compute. Using the Jacobi iterative method as a vehicle, we explore the suitability of stencils on the Grayskull e150. We explore best practice in structuring these codes for the accelerator and demonstrate that the e150 provides similar performance to a Xeon Platinum CPU (albeit BF16 vs FP32) but the e150 uses around five times less energy. Over four e150s we obtain around four times the CPU performance, again at around five times less energy.
format Preprint
id arxiv_https___arxiv_org_abs_2409_18835
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Accelerating stencils on the Tenstorrent Grayskull RISC-V accelerator
Brown, Nick
Barton, Ryan
Distributed, Parallel, and Cluster Computing
The RISC-V Instruction Set Architecture (ISA) has enjoyed phenomenal growth in recent years, however it still to gain popularity in HPC. Whilst adopting RISC-V CPU solutions in HPC might be some way off, RISC-V based PCIe accelerators offer a middle ground where vendors benefit from the flexibility of RISC-V yet fit into existing systems. In this paper we focus on the Tenstorrent Grayskull PCIe RISC-V based accelerator which, built upon Tensix cores, decouples data movement from compute. Using the Jacobi iterative method as a vehicle, we explore the suitability of stencils on the Grayskull e150. We explore best practice in structuring these codes for the accelerator and demonstrate that the e150 provides similar performance to a Xeon Platinum CPU (albeit BF16 vs FP32) but the e150 uses around five times less energy. Over four e150s we obtain around four times the CPU performance, again at around five times less energy.
title Accelerating stencils on the Tenstorrent Grayskull RISC-V accelerator
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2409.18835