Assessing Tenstorrent's RISC-V MatMul Acceleration Capabilities

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cavagna, Hiari Pizzini, Cesarini, Daniele, Bartolini, Andrea
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914307007053824
author Cavagna, Hiari Pizzini
Cesarini, Daniele
Bartolini, Andrea
author_facet Cavagna, Hiari Pizzini
Cesarini, Daniele
Bartolini, Andrea
contents The increasing demand for generative AI as Large Language Models (LLMs) services has driven the need for specialized hardware architectures that optimize computational efficiency and energy consumption. This paper evaluates the performance of the Tenstorrent Grayskull e75 RISC-V accelerator for basic linear algebra kernels at reduced numerical precision, a fundamental operation in LLM computations. We present a detailed characterization of Grayskull's execution model, gridsize, matrix dimensions, data formats, and numerical precision impact computational efficiency. Furthermore, we compare Grayskull's performance against state-of-the-art architectures with tensor acceleration, including Intel Sapphire Rapids processors and two NVIDIA GPUs (V100 and A100). Whilst NVIDIA GPUs dominate raw performance, Grayskull demonstrates a competitive trade-off between power consumption and computational throughput, reaching a peak of 1.55 TFLOPs/Watt with BF16.
format Preprint
id arxiv_https___arxiv_org_abs_2505_06085
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Assessing Tenstorrent's RISC-V MatMul Acceleration Capabilities
Cavagna, Hiari Pizzini
Cesarini, Daniele
Bartolini, Andrea
Performance
Artificial Intelligence
Hardware Architecture
The increasing demand for generative AI as Large Language Models (LLMs) services has driven the need for specialized hardware architectures that optimize computational efficiency and energy consumption. This paper evaluates the performance of the Tenstorrent Grayskull e75 RISC-V accelerator for basic linear algebra kernels at reduced numerical precision, a fundamental operation in LLM computations. We present a detailed characterization of Grayskull's execution model, gridsize, matrix dimensions, data formats, and numerical precision impact computational efficiency. Furthermore, we compare Grayskull's performance against state-of-the-art architectures with tensor acceleration, including Intel Sapphire Rapids processors and two NVIDIA GPUs (V100 and A100). Whilst NVIDIA GPUs dominate raw performance, Grayskull demonstrates a competitive trade-off between power consumption and computational throughput, reaching a peak of 1.55 TFLOPs/Watt with BF16.
title Assessing Tenstorrent's RISC-V MatMul Acceleration Capabilities
topic Performance
Artificial Intelligence
Hardware Architecture
url https://arxiv.org/abs/2505.06085