Shared-PIM: Enabling Concurrent Computation and Data Flow for Faster Processing-in-DRAM

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mamdouh, Ahmed, Geng, Haoran, Niemier, Michael, Hu, Xiaobo Sharon, Reis, Dayane
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909603777740800
author Mamdouh, Ahmed
Geng, Haoran
Niemier, Michael
Hu, Xiaobo Sharon
Reis, Dayane
author_facet Mamdouh, Ahmed
Geng, Haoran
Niemier, Michael
Hu, Xiaobo Sharon
Reis, Dayane
contents Processing-in-Memory (PIM) enhances memory with computational capabilities, potentially solving energy and latency issues associated with data transfer between memory and processors. However, managing concurrent computation and data flow within the PIM architecture incurs significant latency and energy penalty for applications. This paper introduces Shared-PIM, an architecture for in-DRAM PIM that strategically allocates rows in memory banks, bolstered by memory peripherals, for concurrent processing and data movement. Shared-PIM enables simultaneous computation and data transfer within a memory bank. When compared to LISA, a state-of-the-art architecture that facilitates data transfers for in-DRAM PIM, Shared-PIM reduces data movement latency and energy by 5x and 1.2x respectively. Furthermore, when integrated to a state-of-the-art (SOTA) in-DRAM PIM architecture (pLUTo), Shared-PIM achieves 1.4x faster addition and multiplication, and thereby improves the performance of matrix multiplication (MM) tasks by 40%, polynomial multiplication (PMM) by 44%, and numeric number transfer (NTT) tasks by 31%. Moreover, for graph processing tasks like Breadth-First Search (BFS) and Depth-First Search (DFS), Shared-PIM achieves a 29% improvement in speed, all with an area overhead of just 7.16% compared to the baseline pLUTo.
format Preprint
id arxiv_https___arxiv_org_abs_2408_15489
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Shared-PIM: Enabling Concurrent Computation and Data Flow for Faster Processing-in-DRAM
Mamdouh, Ahmed
Geng, Haoran
Niemier, Michael
Hu, Xiaobo Sharon
Reis, Dayane
Hardware Architecture
Processing-in-Memory (PIM) enhances memory with computational capabilities, potentially solving energy and latency issues associated with data transfer between memory and processors. However, managing concurrent computation and data flow within the PIM architecture incurs significant latency and energy penalty for applications. This paper introduces Shared-PIM, an architecture for in-DRAM PIM that strategically allocates rows in memory banks, bolstered by memory peripherals, for concurrent processing and data movement. Shared-PIM enables simultaneous computation and data transfer within a memory bank. When compared to LISA, a state-of-the-art architecture that facilitates data transfers for in-DRAM PIM, Shared-PIM reduces data movement latency and energy by 5x and 1.2x respectively. Furthermore, when integrated to a state-of-the-art (SOTA) in-DRAM PIM architecture (pLUTo), Shared-PIM achieves 1.4x faster addition and multiplication, and thereby improves the performance of matrix multiplication (MM) tasks by 40%, polynomial multiplication (PMM) by 44%, and numeric number transfer (NTT) tasks by 31%. Moreover, for graph processing tasks like Breadth-First Search (BFS) and Depth-First Search (DFS), Shared-PIM achieves a 29% improvement in speed, all with an area overhead of just 7.16% compared to the baseline pLUTo.
title Shared-PIM: Enabling Concurrent Computation and Data Flow for Faster Processing-in-DRAM
topic Hardware Architecture
url https://arxiv.org/abs/2408.15489