Choreographer: A Full-System Framework for Fine-Grained Tasks in Cache Hierarchies

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Hoa, Maidee, Pongstorn, Lowe-Power, Jason, Kaviani, Alireza
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911241389211648
author Nguyen, Hoa
Maidee, Pongstorn
Lowe-Power, Jason
Kaviani, Alireza
author_facet Nguyen, Hoa
Maidee, Pongstorn
Lowe-Power, Jason
Kaviani, Alireza
contents In this paper, we introduce Choreographer, a simulation framework that enables a holistic system-level evaluation of fine-grained accelerators designed for latency-sensitive tasks. Unlike existing frameworks, Choreographer captures all hardware and software overheads in core-accelerator and cache-accelerator interactions, integrating a detailed gem5-based hardware stack featuring an AMBA coherent hub interface (CHI) mesh network and a complete Linux-based software stack. To facilitate rapid prototyping, it offers a C++ application programming interface and modular configuration options. Our detailed cache model provides accurate insights into performance variations caused by cache configurations, which are not captured by other frameworks. The framework is demonstrated through two case studies: a data-aware prefetcher for graph analytics workloads, and a quicksort accelerator. Our evaluation shows that the prefetcher achieves speedups between 1.08x and 1.88x by reducing memory access latency, while the quicksort accelerator delivers more than 2x speedup with minimal address translation overhead. These findings underscore the ability of Choreographer to model complex hardware-software interactions and optimize performance in small task offloading scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2510_26944
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Choreographer: A Full-System Framework for Fine-Grained Tasks in Cache Hierarchies
Nguyen, Hoa
Maidee, Pongstorn
Lowe-Power, Jason
Kaviani, Alireza
Hardware Architecture
In this paper, we introduce Choreographer, a simulation framework that enables a holistic system-level evaluation of fine-grained accelerators designed for latency-sensitive tasks. Unlike existing frameworks, Choreographer captures all hardware and software overheads in core-accelerator and cache-accelerator interactions, integrating a detailed gem5-based hardware stack featuring an AMBA coherent hub interface (CHI) mesh network and a complete Linux-based software stack. To facilitate rapid prototyping, it offers a C++ application programming interface and modular configuration options. Our detailed cache model provides accurate insights into performance variations caused by cache configurations, which are not captured by other frameworks. The framework is demonstrated through two case studies: a data-aware prefetcher for graph analytics workloads, and a quicksort accelerator. Our evaluation shows that the prefetcher achieves speedups between 1.08x and 1.88x by reducing memory access latency, while the quicksort accelerator delivers more than 2x speedup with minimal address translation overhead. These findings underscore the ability of Choreographer to model complex hardware-software interactions and optimize performance in small task offloading scenarios.
title Choreographer: A Full-System Framework for Fine-Grained Tasks in Cache Hierarchies
topic Hardware Architecture
url https://arxiv.org/abs/2510.26944