AME-PIM: Can Memory be Your Next Tensor Accelerator?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Venieri, Emanuele, Manoni, Simone, Florian, Alberto, Park, Jaehyun, Sohn, Kyomin, Bartolini, Andrea
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913076891090944
author Venieri, Emanuele
Manoni, Simone
Florian, Alberto
Park, Jaehyun
Sohn, Kyomin
Bartolini, Andrea
author_facet Venieri, Emanuele
Manoni, Simone
Florian, Alberto
Park, Jaehyun
Sohn, Kyomin
Bartolini, Andrea
contents High Bandwidth Memory with Processing-in-Memory (HBM-PIM) offers an opportunity to reduce data movement by executing computation directly inside memory, but current commercial platforms expose limited instruction sets and require specialized software stacks. In this work, we investigate whether HBM-PIM can serve as a backend for ISA-level matrix acceleration, using the RISC-V Attached Matrix Extension (AME) as a semantic reference. We propose a PEP-based execution model that maps AME element-wise and matrix instructions to HBM-PIM micro-kernels and data instructions in memory operations. Differently from SoA HBM-PIM, we introduce a reduction-free outer-product dataflow that enables accumulation entirely within memory despite the lack of native reduction support. Our approach supports end-to-end execution of element-wise operations, GEMV, and GEMM in PIM mode, minimizing host involvement and off-chip transfers. An experimental evaluation on Samsung Aquabolt-XL shows that AME matrix tile multiplication achieves up to 14.9 GFLOP/s (59.4 FLOP/cycle) on a single HBM pseudo-channel.
format Preprint
id arxiv_https___arxiv_org_abs_2604_27808
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AME-PIM: Can Memory be Your Next Tensor Accelerator?
Venieri, Emanuele
Manoni, Simone
Florian, Alberto
Park, Jaehyun
Sohn, Kyomin
Bartolini, Andrea
Hardware Architecture
High Bandwidth Memory with Processing-in-Memory (HBM-PIM) offers an opportunity to reduce data movement by executing computation directly inside memory, but current commercial platforms expose limited instruction sets and require specialized software stacks. In this work, we investigate whether HBM-PIM can serve as a backend for ISA-level matrix acceleration, using the RISC-V Attached Matrix Extension (AME) as a semantic reference. We propose a PEP-based execution model that maps AME element-wise and matrix instructions to HBM-PIM micro-kernels and data instructions in memory operations. Differently from SoA HBM-PIM, we introduce a reduction-free outer-product dataflow that enables accumulation entirely within memory despite the lack of native reduction support. Our approach supports end-to-end execution of element-wise operations, GEMV, and GEMM in PIM mode, minimizing host involvement and off-chip transfers. An experimental evaluation on Samsung Aquabolt-XL shows that AME matrix tile multiplication achieves up to 14.9 GFLOP/s (59.4 FLOP/cycle) on a single HBM pseudo-channel.
title AME-PIM: Can Memory be Your Next Tensor Accelerator?
topic Hardware Architecture
url https://arxiv.org/abs/2604.27808