The Configuration Wall: Characterization and Elimination of Accelerator Configuration Overhead
Fuente:
arXiv
Saved in:
| Main Authors: | Van Delm, Josse, Lydike, Anton, Dumoulin, Joren, Crols, Jonas, Yi, Xiaoling, Antonio, Ryan, Woodruff, Jackson, Grosser, Tobias, Verhelst, Marian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Stream parallel skeleton optimization
by: Aldinucci, Marco, et al.
Published: (2024)
by: Aldinucci, Marco, et al.
Published: (2024)
StreamFlow: cross-breeding cloud with HPC
by: Colonnelli, Iacopo, et al.
Published: (2020)
by: Colonnelli, Iacopo, et al.
Published: (2020)
Hybrid Quantum-HPC Middleware Systems for Adaptive Resource, Workload and Task Management
by: Mantha, Pradeep, et al.
Published: (2026)
by: Mantha, Pradeep, et al.
Published: (2026)
Janus: Compiler-Based Defense Against Transient Execution Attacks Using ARM Hardware Primitives
by: Ouyang, Ciyan, et al.
Published: (2026)
by: Ouyang, Ciyan, et al.
Published: (2026)
Kernel Looping: Eliminating Synchronization Boundaries for Peak Inference Performance
by: Koeplinger, David, et al.
Published: (2024)
by: Koeplinger, David, et al.
Published: (2024)
A Multi-level Compiler Backend for Accelerated Micro-kernels Targeting RISC-V ISA Extensions
by: Lopoukhine, Alexandre, et al.
Published: (2025)
by: Lopoukhine, Alexandre, et al.
Published: (2025)
A Compilation Framework for Quantum Circuits with Mid-Circuit Measurement Error Awareness
by: Zhong, Ming, et al.
Published: (2025)
by: Zhong, Ming, et al.
Published: (2025)
PoCL-R: An Open Standard Based Offloading Layer for Heterogeneous Multi-Access Edge Computing with Server Side Scalability
by: Solanti, Jan, et al.
Published: (2023)
by: Solanti, Jan, et al.
Published: (2023)
Mapping Sparse Triangular Solves to GPUs via Fine-grained Domain Decomposition
by: Gondhalekar, Atharva, et al.
Published: (2025)
by: Gondhalekar, Atharva, et al.
Published: (2025)
SpaDA: A Spatial Dataflow Architecture Programming Language
by: Gianinazzi, Lukas, et al.
Published: (2025)
by: Gianinazzi, Lukas, et al.
Published: (2025)
Conversational Concurrency
by: Garnock-Jones, Tony
Published: (2024)
by: Garnock-Jones, Tony
Published: (2024)
Ember: A Compiler for Efficient Embedding Operations on Decoupled Access-Execute Architectures
by: Siracusa, Marco, et al.
Published: (2025)
by: Siracusa, Marco, et al.
Published: (2025)
Assembly of FETI dual operator using CUDA
by: Homola, Jakub, et al.
Published: (2025)
by: Homola, Jakub, et al.
Published: (2025)
Scheduler-Driven Job Atomization
by: Konopa, Michal, et al.
Published: (2025)
by: Konopa, Michal, et al.
Published: (2025)
JASDA: Introducing Job-Aware Scheduling in Scheduler-Driven Job Atomization
by: Konopa, Michal, et al.
Published: (2025)
by: Konopa, Michal, et al.
Published: (2025)
MATCH: Model-Aware TVM-based Compilation for Heterogeneous Edge Devices
by: Hamdi, Mohamed Amine, et al.
Published: (2024)
by: Hamdi, Mohamed Amine, et al.
Published: (2024)
NVLang: Unified Static Typing for Actor-Based Concurrency on the BEAM
by: Guerreiro, Miguel de Oliveira
Published: (2025)
by: Guerreiro, Miguel de Oliveira
Published: (2025)
Fancy Some Chips for Your TeaStore? Modeling the Control of an Adaptable Discrete System
by: Gallone, Anna, et al.
Published: (2025)
by: Gallone, Anna, et al.
Published: (2025)
Dormancy-aware timed branching bisimilarity, with an application to communication protocol analysis
by: Middelburg, C. A.
Published: (2021)
by: Middelburg, C. A.
Published: (2021)
NM-SpMM: Accelerating Matrix Multiplication Using N:M Sparsity with GPGPU
by: Ma, Cong, et al.
Published: (2025)
by: Ma, Cong, et al.
Published: (2025)
SynapticCore-X: A Modular Neural Processing Architecture for Low-Cost FPGA Acceleration
by: Parameshwara, Arya
Published: (2025)
by: Parameshwara, Arya
Published: (2025)
PIM-STM: Software Transactional Memory for Processing-In-Memory Systems
by: Lopes, André, et al.
Published: (2024)
by: Lopes, André, et al.
Published: (2024)
MEDEA: A Design-Time Multi-Objective Manager for Energy-Efficient DNN Inference on Heterogeneous Ultra-Low Power Platforms
by: Taji, Hossein, et al.
Published: (2025)
by: Taji, Hossein, et al.
Published: (2025)
Securing Mixed Rust with Hardware Capabilities
by: Yu, Jason Zhijingcheng, et al.
Published: (2025)
by: Yu, Jason Zhijingcheng, et al.
Published: (2025)
GraphPerf-RT: A Graph-Driven Performance Model for Hardware-Aware Scheduling of OpenMP Codes
by: Pivezhandi, Mohammad, et al.
Published: (2025)
by: Pivezhandi, Mohammad, et al.
Published: (2025)
Rhizomes and Diffusions for Processing Highly Skewed Graphs on Fine-Grain Message-Driven Systems
by: Chandio, Bibrak Qamar, et al.
Published: (2024)
by: Chandio, Bibrak Qamar, et al.
Published: (2024)
Reasoning about distributive laws in a concurrent refinement algebra
by: Meinicke, Larissa A., et al.
Published: (2024)
by: Meinicke, Larissa A., et al.
Published: (2024)
Restructuring a concurrent refinement algebra
by: Hayes, Ian J., et al.
Published: (2024)
by: Hayes, Ian J., et al.
Published: (2024)
Modelling Distributed Applications with Mixed-Choice Stateful Typestates
by: Parrinha, Francisco, et al.
Published: (2026)
by: Parrinha, Francisco, et al.
Published: (2026)
FlexiBit: Fully Flexible Precision Bit-parallel Accelerator Architecture for Arbitrary Mixed Precision AI
by: Tahmasebi, Faraz, et al.
Published: (2024)
by: Tahmasebi, Faraz, et al.
Published: (2024)
OpenEye: A Scalable Open-Source Hardware Accelerator for DNNs
by: Lebold, Denis, et al.
Published: (2026)
by: Lebold, Denis, et al.
Published: (2026)
A Gentle Overview of Asynchronous Session-based Concurrency: Deadlock Freedom by Typing
by: Heuvel, Bas van den, et al.
Published: (2024)
by: Heuvel, Bas van den, et al.
Published: (2024)
Coflex: Enhancing HW-NAS with Sparse Gaussian Processes for Efficient and Scalable DNN Accelerator Design
by: Ma, Yinhui, et al.
Published: (2025)
by: Ma, Yinhui, et al.
Published: (2025)
Accelerating cosmological simulations on GPUs: a portable approach using OpenMP
by: Lepinzan, M. D., et al.
Published: (2025)
by: Lepinzan, M. D., et al.
Published: (2025)
Dynamic Race Detection With O(1) Samples
by: Thokair, Mosaad Al, et al.
Published: (2025)
by: Thokair, Mosaad Al, et al.
Published: (2025)
The B2Scala Tool: Integrating Bach in Scala with Security in Mind
by: Ouardi, Doha, et al.
Published: (2024)
by: Ouardi, Doha, et al.
Published: (2024)
ASTER: Attention-based Spiking Transformer Engine for Event-driven Reasoning
by: Das, Tamoghno, et al.
Published: (2025)
by: Das, Tamoghno, et al.
Published: (2025)
Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories
by: Lee, Ming-Yen, et al.
Published: (2025)
by: Lee, Ming-Yen, et al.
Published: (2025)
Exploring the Design Space for Message-Driven Systems for Dynamic Graph Processing using CCA
by: Chandio, Bibrak Qamar, et al.
Published: (2024)
by: Chandio, Bibrak Qamar, et al.
Published: (2024)
Utilizing Sparsity in the GPU-accelerated Assembly of Schur Complement Matrices in Domain Decomposition Methods
by: Homola, Jakub, et al.
Published: (2025)
by: Homola, Jakub, et al.
Published: (2025)
Similar Items
-
Stream parallel skeleton optimization
by: Aldinucci, Marco, et al.
Published: (2024) -
StreamFlow: cross-breeding cloud with HPC
by: Colonnelli, Iacopo, et al.
Published: (2020) -
Hybrid Quantum-HPC Middleware Systems for Adaptive Resource, Workload and Task Management
by: Mantha, Pradeep, et al.
Published: (2026) -
Janus: Compiler-Based Defense Against Transient Execution Attacks Using ARM Hardware Primitives
by: Ouyang, Ciyan, et al.
Published: (2026) -
Kernel Looping: Eliminating Synchronization Boundaries for Peak Inference Performance
by: Koeplinger, David, et al.
Published: (2024)