Proteus: Enabling High-Performance Processing-Using-DRAM with Dynamic Bit-Precision, Adaptive Data Representation, and Flexible Arithmetic

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Oliveira, Geraldo F., Kabra, Mayank, Guo, Yuxin, Chen, Kangqi, Yağlıkçı, A. Giray, Soysal, Melina, Sadrosadati, Mohammad, Bueno, Joaquin Olivares, Ghose, Saugata, Gómez-Luna, Juan, Mutlu, Onur
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916791525048320
author Oliveira, Geraldo F.
Kabra, Mayank
Guo, Yuxin
Chen, Kangqi
Yağlıkçı, A. Giray
Soysal, Melina
Sadrosadati, Mohammad
Bueno, Joaquin Olivares
Ghose, Saugata
Gómez-Luna, Juan
Mutlu, Onur
author_facet Oliveira, Geraldo F.
Kabra, Mayank
Guo, Yuxin
Chen, Kangqi
Yağlıkçı, A. Giray
Soysal, Melina
Sadrosadati, Mohammad
Bueno, Joaquin Olivares
Ghose, Saugata
Gómez-Luna, Juan
Mutlu, Onur
contents Processing-using-DRAM (PUD) is a paradigm where the analog operational properties of DRAM are used to perform bulk logic operations. While PUD promises high throughput at low energy and area cost, we uncover three limitations of existing PUD approaches that lead to significant inefficiencies: (i) static data representation, i.e., two's complement with fixed bit-precision, leading to unnecessary computation over useless (i.e., inconsequential) data; (ii) support for only throughput-oriented execution, where the high latency of individual PUD operations can only be hidden in the presence of bulk data-level parallelism; and (iii) high latency for high-precision (e.g., 32-bit) operations. To address these issues, we propose Proteus, the first hardware framework that addresses the high execution latency of bulk bitwise PUD operations by implementing a data-aware runtime engine for PUD. Proteus reduces the latency of PUD operations in three different ways: (i) Proteus dynamically reduces the bit-precision (and thus the latency and energy consumption) of PUD operations by exploiting narrow values (i.e., values with many leading zeros or ones); (ii) Proteus concurrently executes independent in-DRAM primitives belonging to a single PUD operation across multiple DRAM arrays; (iii) Proteus chooses and uses the most appropriate data representation and arithmetic algorithm implementation for a given PUD instruction transparently to the programmer.
format Preprint
id arxiv_https___arxiv_org_abs_2501_17466
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Proteus: Enabling High-Performance Processing-Using-DRAM with Dynamic Bit-Precision, Adaptive Data Representation, and Flexible Arithmetic
Oliveira, Geraldo F.
Kabra, Mayank
Guo, Yuxin
Chen, Kangqi
Yağlıkçı, A. Giray
Soysal, Melina
Sadrosadati, Mohammad
Bueno, Joaquin Olivares
Ghose, Saugata
Gómez-Luna, Juan
Mutlu, Onur
Hardware Architecture
Distributed, Parallel, and Cluster Computing
Processing-using-DRAM (PUD) is a paradigm where the analog operational properties of DRAM are used to perform bulk logic operations. While PUD promises high throughput at low energy and area cost, we uncover three limitations of existing PUD approaches that lead to significant inefficiencies: (i) static data representation, i.e., two's complement with fixed bit-precision, leading to unnecessary computation over useless (i.e., inconsequential) data; (ii) support for only throughput-oriented execution, where the high latency of individual PUD operations can only be hidden in the presence of bulk data-level parallelism; and (iii) high latency for high-precision (e.g., 32-bit) operations. To address these issues, we propose Proteus, the first hardware framework that addresses the high execution latency of bulk bitwise PUD operations by implementing a data-aware runtime engine for PUD. Proteus reduces the latency of PUD operations in three different ways: (i) Proteus dynamically reduces the bit-precision (and thus the latency and energy consumption) of PUD operations by exploiting narrow values (i.e., values with many leading zeros or ones); (ii) Proteus concurrently executes independent in-DRAM primitives belonging to a single PUD operation across multiple DRAM arrays; (iii) Proteus chooses and uses the most appropriate data representation and arithmetic algorithm implementation for a given PUD instruction transparently to the programmer.
title Proteus: Enabling High-Performance Processing-Using-DRAM with Dynamic Bit-Precision, Adaptive Data Representation, and Flexible Arithmetic
topic Hardware Architecture
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2501.17466