Performance Evaluation of a Next-Generation SX-Aurora TSUBASA Vector Supercomputer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Takahashi, Keichi, Fujimoto, Soya, Nagase, Satoru, Isobe, Yoko, Shimomura, Yoichi, Egawa, Ryusuke, Takizawa, Hiroyuki
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909286814187520
author Takahashi, Keichi
Fujimoto, Soya
Nagase, Satoru
Isobe, Yoko
Shimomura, Yoichi
Egawa, Ryusuke
Takizawa, Hiroyuki
author_facet Takahashi, Keichi
Fujimoto, Soya
Nagase, Satoru
Isobe, Yoko
Shimomura, Yoichi
Egawa, Ryusuke
Takizawa, Hiroyuki
contents Data movement is a key bottleneck in terms of both performance and energy efficiency in modern HPC systems. The NEC SX-series supercomputers have a long history of accelerating memory-intensive HPC applications by providing sufficient memory bandwidth to applications. In this paper, we analyze the performance of a prototype SX-Aurora TSUBASA supercomputer equipped with the brand-new Vector Engine (VE30) processor. VE30 is the first major update to the Vector Engine processor series, and offers significantly improved memory access performance due to its renewed memory subsystem. Moreover, it introduces new instructions and incorporates architectural advancements tailored for accelerating memory-intensive applications. Using standard benchmarks, we demonstrate that VE30 considerably outperforms other processors in both performance and efficiency of memory-intensive applications. We also evaluate VE30 using applications including SPEChpc, and show that VE30 can run real-world applications with high performance. Finally, we discuss performance tuning techniques to obtain maximum performance from VE30.
format Preprint
id arxiv_https___arxiv_org_abs_2304_11921
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Performance Evaluation of a Next-Generation SX-Aurora TSUBASA Vector Supercomputer
Takahashi, Keichi
Fujimoto, Soya
Nagase, Satoru
Isobe, Yoko
Shimomura, Yoichi
Egawa, Ryusuke
Takizawa, Hiroyuki
Distributed, Parallel, and Cluster Computing
Performance
Data movement is a key bottleneck in terms of both performance and energy efficiency in modern HPC systems. The NEC SX-series supercomputers have a long history of accelerating memory-intensive HPC applications by providing sufficient memory bandwidth to applications. In this paper, we analyze the performance of a prototype SX-Aurora TSUBASA supercomputer equipped with the brand-new Vector Engine (VE30) processor. VE30 is the first major update to the Vector Engine processor series, and offers significantly improved memory access performance due to its renewed memory subsystem. Moreover, it introduces new instructions and incorporates architectural advancements tailored for accelerating memory-intensive applications. Using standard benchmarks, we demonstrate that VE30 considerably outperforms other processors in both performance and efficiency of memory-intensive applications. We also evaluate VE30 using applications including SPEChpc, and show that VE30 can run real-world applications with high performance. Finally, we discuss performance tuning techniques to obtain maximum performance from VE30.
title Performance Evaluation of a Next-Generation SX-Aurora TSUBASA Vector Supercomputer
topic Distributed, Parallel, and Cluster Computing
Performance
url https://arxiv.org/abs/2304.11921