HERMES: High-Performance RISC-V Memory Hierarchy for ML Workloads

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Suryadevara, Pranav
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915211906121728
author Suryadevara, Pranav
author_facet Suryadevara, Pranav
contents The growth of machine learning (ML) workloads has underscored the importance of efficient memory hierarchies to address bandwidth, latency, and scalability challenges. HERMES focuses on optimizing memory subsystems for RISC-V architectures to meet the computational needs of ML models such as CNNs, RNNs, and Transformers. This project explores state-of-the-art techniques such as advanced prefetching, tensor-aware caching, and hybrid memory models. The cornerstone of HERMES is the integration of shared L3 caches with fine-grained coherence protocols equipped with specialized pathways to deep-learning accelerators such as Gemmini. Simulation tools like Gem5 and DRAMSim2 were used to evaluate baseline performance and scalability under representative ML workloads. The findings of this study highlight the design choices, and the anticipated challenges, paving the way for low-latency scalable memory operations for ML applications.
format Preprint
id arxiv_https___arxiv_org_abs_2503_13064
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HERMES: High-Performance RISC-V Memory Hierarchy for ML Workloads
Suryadevara, Pranav
Hardware Architecture
Performance
B.3.2; C.1.3; C.3
The growth of machine learning (ML) workloads has underscored the importance of efficient memory hierarchies to address bandwidth, latency, and scalability challenges. HERMES focuses on optimizing memory subsystems for RISC-V architectures to meet the computational needs of ML models such as CNNs, RNNs, and Transformers. This project explores state-of-the-art techniques such as advanced prefetching, tensor-aware caching, and hybrid memory models. The cornerstone of HERMES is the integration of shared L3 caches with fine-grained coherence protocols equipped with specialized pathways to deep-learning accelerators such as Gemmini. Simulation tools like Gem5 and DRAMSim2 were used to evaluate baseline performance and scalability under representative ML workloads. The findings of this study highlight the design choices, and the anticipated challenges, paving the way for low-latency scalable memory operations for ML applications.
title HERMES: High-Performance RISC-V Memory Hierarchy for ML Workloads
topic Hardware Architecture
Performance
B.3.2; C.1.3; C.3
url https://arxiv.org/abs/2503.13064