High-performance Vector-length Agnostic Quantum Circuit Simulations on ARM Processors

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shi, Ruimin, Schieffer, Gabin, Lin, Pei-Hung, Gokhale, Maya, Herten, Andreas, Peng, Ivy
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917336052662272
author Shi, Ruimin
Schieffer, Gabin
Lin, Pei-Hung
Gokhale, Maya
Herten, Andreas
Peng, Ivy
author_facet Shi, Ruimin
Schieffer, Gabin
Lin, Pei-Hung
Gokhale, Maya
Herten, Andreas
Peng, Ivy
contents ARM SVE and RISC-V RVV are emerging vector architectures in high-end processors that support vectorization of flexible vector length. In this work, we leverage an important workload for quantum computing, quantum state-vector simulations, to understand whether high-performance portability can be achieved in a vector-length agnostic (VLA) design. We propose a VLA design and optimization techniques critical for achieving high performance, including VLEN-adaptive memory layout adjustment, load buffering, fine-grained loop control, and gate fusion-based arithmetic intensity adaptation. We provide an implementation in Google's Qsim and evaluate five quantum circuits of up to 36 qubits on three ARM processors, including NVIDIA Grace, AWS Graviton3, and Fujitsu A64FX. By defining new metrics and PMU events to quantify vectorization activities, we draw generic insights for future VLA designs. Our single-source implementation of VLA quantum simulations achieves up to 4.5x speedup on A64FX, 2.5x speedup on Grace, and 1.5x speedup on Graviton.
format Preprint
id arxiv_https___arxiv_org_abs_2602_09604
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle High-performance Vector-length Agnostic Quantum Circuit Simulations on ARM Processors
Shi, Ruimin
Schieffer, Gabin
Lin, Pei-Hung
Gokhale, Maya
Herten, Andreas
Peng, Ivy
Distributed, Parallel, and Cluster Computing
ARM SVE and RISC-V RVV are emerging vector architectures in high-end processors that support vectorization of flexible vector length. In this work, we leverage an important workload for quantum computing, quantum state-vector simulations, to understand whether high-performance portability can be achieved in a vector-length agnostic (VLA) design. We propose a VLA design and optimization techniques critical for achieving high performance, including VLEN-adaptive memory layout adjustment, load buffering, fine-grained loop control, and gate fusion-based arithmetic intensity adaptation. We provide an implementation in Google's Qsim and evaluate five quantum circuits of up to 36 qubits on three ARM processors, including NVIDIA Grace, AWS Graviton3, and Fujitsu A64FX. By defining new metrics and PMU events to quantify vectorization activities, we draw generic insights for future VLA designs. Our single-source implementation of VLA quantum simulations achieves up to 4.5x speedup on A64FX, 2.5x speedup on Grace, and 1.5x speedup on Graviton.
title High-performance Vector-length Agnostic Quantum Circuit Simulations on ARM Processors
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2602.09604