Closer in the Gap: Towards Portable Performance on RISC-V Vector Processors

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shi, Ruimin, Gokhale, Maya, Lin, Pei-Hung, Teruel, Xavier, Peng, Ivy
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911706647625728
author Shi, Ruimin
Gokhale, Maya
Lin, Pei-Hung
Teruel, Xavier
Peng, Ivy
author_facet Shi, Ruimin
Gokhale, Maya
Lin, Pei-Hung
Teruel, Xavier
Peng, Ivy
contents The RISC-V Vector Extension~(RVV) is a cornerstone for supporting compute throughout in scientific and machine learning workloads. Yet compiler support and performance monitoring on real RVV~1.0 hardware are still evolving. In this work, we design a suite of assembly microbenchmarks to establish performance ceilings and calibrate performance counters on RVV hardware. Leveraging the assembly benchmarks, we find that predication overhead and stride load pose performance challenges that current compiler cost models do not yet fully address. Moreover, we present the first evaluation of GCC~15 and LLVM~21 autovectorization in HPC and ML proxy applications. GCC~15 outperforms LLVM~21 in four out of six applications. LLVM~21 only outperforms GCC~15 in SGEMM and DGEMM, driven by more aggressive instruction reduction confirmed through validated \texttt{perf} counters on the RVV hardware. We further show that the default LMUL selection in compilers performs close to the optimal. To study the RVV support for product-level application, we also evaluate the state-vector quantum simulator, Google's Qsim, with both manual RVV intrinsics and compiler auto-vectorization, revealing immaturity in current RVV compiler for complicated memory access pattern.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10860
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Closer in the Gap: Towards Portable Performance on RISC-V Vector Processors
Shi, Ruimin
Gokhale, Maya
Lin, Pei-Hung
Teruel, Xavier
Peng, Ivy
Distributed, Parallel, and Cluster Computing
The RISC-V Vector Extension~(RVV) is a cornerstone for supporting compute throughout in scientific and machine learning workloads. Yet compiler support and performance monitoring on real RVV~1.0 hardware are still evolving. In this work, we design a suite of assembly microbenchmarks to establish performance ceilings and calibrate performance counters on RVV hardware. Leveraging the assembly benchmarks, we find that predication overhead and stride load pose performance challenges that current compiler cost models do not yet fully address. Moreover, we present the first evaluation of GCC~15 and LLVM~21 autovectorization in HPC and ML proxy applications. GCC~15 outperforms LLVM~21 in four out of six applications. LLVM~21 only outperforms GCC~15 in SGEMM and DGEMM, driven by more aggressive instruction reduction confirmed through validated \texttt{perf} counters on the RVV hardware. We further show that the default LMUL selection in compilers performs close to the optimal. To study the RVV support for product-level application, we also evaluate the state-vector quantum simulator, Google's Qsim, with both manual RVV intrinsics and compiler auto-vectorization, revealing immaturity in current RVV compiler for complicated memory access pattern.
title Closer in the Gap: Towards Portable Performance on RISC-V Vector Processors
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2605.10860