ACS: Concurrent Kernel Execution on Irregular, Input-Dependent Computational Graphs
Fuente:
arXiv
Saved in:
| Main Authors: | Durvasula, Sankeerth, Zhao, Adrian, Kiguru, Raymond, Guan, Yushi, Chen, Zhonghan, Vijaykumar, Nandita |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FG-Attn: Leveraging Fine-Grained Sparsity In Diffusion Transformers
by: Durvasula, Sankeerth, et al.
Published: (2025)
by: Durvasula, Sankeerth, et al.
Published: (2025)
DCC: Data-Centric Compilation of Machine Learning Kernels for Processing-In-Memory Architectures
by: Yang, Peiming, et al.
Published: (2025)
by: Yang, Peiming, et al.
Published: (2025)
Shared-PIM: Enabling Concurrent Computation and Data Flow for Faster Processing-in-DRAM
by: Mamdouh, Ahmed, et al.
Published: (2024)
by: Mamdouh, Ahmed, et al.
Published: (2024)
Efficient Sparse Processing-in-Memory Architecture (ESPIM) for Machine Learning Inference
by: He, Mingxuan, et al.
Published: (2024)
by: He, Mingxuan, et al.
Published: (2024)
Squire: A General-Purpose Accelerator to Exploit Fine-Grain Parallelism on Dependency-Bound Kernels
by: Langarita, Rubén, et al.
Published: (2025)
by: Langarita, Rubén, et al.
Published: (2025)
QED: Scalable Verification of Hardware Memory Consistency
by: Ravi, Gokulan, et al.
Published: (2024)
by: Ravi, Gokulan, et al.
Published: (2024)
SLTarch: Towards Scalable Point-Based Neural Rendering by Taming Workload Imbalance and Memory Irregularity
by: Li, Xingyang, et al.
Published: (2025)
by: Li, Xingyang, et al.
Published: (2025)
Hardware-Efficient CNNs: Interleaved Approximate FP32 Multipliers for Kernel Computation
by: Gowda, Bindu G, et al.
Published: (2025)
by: Gowda, Bindu G, et al.
Published: (2025)
Data-aware Dynamic Execution of Irregular Workloads on Heterogeneous Systems
by: Bai, Zhenyu, et al.
Published: (2025)
by: Bai, Zhenyu, et al.
Published: (2025)
ICP: Exploiting Instruction Correlation for Prefetching Irregular Memory Accesses
by: Li, Mengming, et al.
Published: (2026)
by: Li, Mengming, et al.
Published: (2026)
PyGim: An Efficient Graph Neural Network Library for Real Processing-In-Memory Architectures
by: Giannoula, Christina, et al.
Published: (2024)
by: Giannoula, Christina, et al.
Published: (2024)
Nexus Machine: An Active Message Inspired Reconfigurable Architecture for Irregular Workloads
by: Juneja, Rohan, et al.
Published: (2025)
by: Juneja, Rohan, et al.
Published: (2025)
UpDown: Programmable fine-grained Events for Scalable Performance on Irregular Applications
by: Rajasukumar, Andronicus, et al.
Published: (2024)
by: Rajasukumar, Andronicus, et al.
Published: (2024)
DARE: An Irregularity-Tolerant Matrix Processing Unit with a Densifying ISA and Filtered Runahead Execution
by: Yang, Xin, et al.
Published: (2025)
by: Yang, Xin, et al.
Published: (2025)
NEURAL: An Elastic Neuromorphic Architecture with Hybrid Data-Event Execution and On-the-fly Attention Dataflow
by: Chen, Yuehai, et al.
Published: (2025)
by: Chen, Yuehai, et al.
Published: (2025)
Double Duty: FPGA Architecture to Enable Concurrent LUT and Adder Chain Usage
by: Pun, Junius, et al.
Published: (2025)
by: Pun, Junius, et al.
Published: (2025)
A Direct Memory Access Controller (DMAC) for Irregular Data Transfers on RISC-V Linux Systems
by: Benz, Thomas, et al.
Published: (2025)
by: Benz, Thomas, et al.
Published: (2025)
Efficient and Accurate Graph Classification with Hyperdimensional Computing on FPGA
by: Arockiaraj, Jebacyril, et al.
Published: (2025)
by: Arockiaraj, Jebacyril, et al.
Published: (2025)
Compilation and Execution of an Embeddable YOLO-NAS on the VTA
by: Faure-Gignoux, Anthony, et al.
Published: (2026)
by: Faure-Gignoux, Anthony, et al.
Published: (2026)
Kernel Approximation using Analog In-Memory Computing
by: Büchel, Julian, et al.
Published: (2024)
by: Büchel, Julian, et al.
Published: (2024)
Virtuoso: Enabling Fast and Accurate Virtual Memory Research via an Imitation-based Operating System Simulation Methodology
by: Kanellopoulos, Konstantinos, et al.
Published: (2024)
by: Kanellopoulos, Konstantinos, et al.
Published: (2024)
A Data-Driven Dynamic Execution Orchestration Architecture
by: Bai, Zhenyu, et al.
Published: (2026)
by: Bai, Zhenyu, et al.
Published: (2026)
Affordable HPC: Leveraging Small Clusters for Big Data and Graph Computing
by: Wu, Ruilong, et al.
Published: (2024)
by: Wu, Ruilong, et al.
Published: (2024)
A Hybrid Delay Model for Interconnected Multi-Input Gates
by: Ferdowsi, Arman, et al.
Published: (2024)
by: Ferdowsi, Arman, et al.
Published: (2024)
PIVOT- Input-aware Path Selection for Energy-efficient ViT Inference
by: Moitra, Abhishek, et al.
Published: (2024)
by: Moitra, Abhishek, et al.
Published: (2024)
Optimising GPGPU Execution Through Runtime Micro-Architecture Parameter Analysis
by: Sarda, Giuseppe M., et al.
Published: (2024)
by: Sarda, Giuseppe M., et al.
Published: (2024)
Mapping and Execution of Nested Loops on Processor Arrays: CGRAs vs. TCPAs
by: Walter, Dominik, et al.
Published: (2025)
by: Walter, Dominik, et al.
Published: (2025)
ERASER: Efficient RTL FAult Simulation Framework with Trimmed Execution Redundancy
by: Tang, Jiaping, et al.
Published: (2025)
by: Tang, Jiaping, et al.
Published: (2025)
Accelerating Mini-batch HGNN Training by Reducing CUDA Kernels
by: Wu, Meng, et al.
Published: (2024)
by: Wu, Meng, et al.
Published: (2024)
LLM-Powered Code Analysis and Optimization for Gaussian Splatting Kernels
by: Hu, Yi, et al.
Published: (2025)
by: Hu, Yi, et al.
Published: (2025)
Constable: Improving Performance and Power Efficiency by Safely Eliminating Load Instruction Execution
by: Bera, Rahul, et al.
Published: (2024)
by: Bera, Rahul, et al.
Published: (2024)
A System Architecture for Low Latency Multiprogramming Quantum Computing
by: Zhao, Yilun, et al.
Published: (2026)
by: Zhao, Yilun, et al.
Published: (2026)
LLMulator: Generalizable Cost Modeling for Dataflow Accelerators with Input-Adaptive Control Flow
by: Chang, Kaiyan, et al.
Published: (2025)
by: Chang, Kaiyan, et al.
Published: (2025)
SkipOPU: An FPGA-based Overlay Processor for Large Language Models with Dynamically Allocated Computation
by: He, Zicheng, et al.
Published: (2026)
by: He, Zicheng, et al.
Published: (2026)
Lit Silicon: A Case Where Thermal Imbalance Couples Concurrent Execution in Multiple GPUs
by: Kurzynski, Marco, et al.
Published: (2025)
by: Kurzynski, Marco, et al.
Published: (2025)
Design and Optimization of Mixed-Kernel Mixed-Signal SVMs for Flexible Electronics
by: Afentaki, Florentia, et al.
Published: (2025)
by: Afentaki, Florentia, et al.
Published: (2025)
Decoupled Access-Execute enabled DVFS for tinyML deployments on STM32 microcontrollers
by: Alvanaki, Elisavet Lydia, et al.
Published: (2024)
by: Alvanaki, Elisavet Lydia, et al.
Published: (2024)
Late Breaking Results: Boosting Efficient Dual-Issue Execution on Lightweight RISC-V Cores
by: Colagrande, Luca, et al.
Published: (2026)
by: Colagrande, Luca, et al.
Published: (2026)
RealProbe: An Automated and Lightweight Performance Profiler for In-FPGA Execution of High-Level Synthesis Designs
by: Kim, Jiho, et al.
Published: (2025)
by: Kim, Jiho, et al.
Published: (2025)
Efficient Kernel Mapping and Comprehensive System Evaluation of LLM Acceleration on a CGLA
by: Ando, Takuto, et al.
Published: (2025)
by: Ando, Takuto, et al.
Published: (2025)
Similar Items
-
FG-Attn: Leveraging Fine-Grained Sparsity In Diffusion Transformers
by: Durvasula, Sankeerth, et al.
Published: (2025) -
DCC: Data-Centric Compilation of Machine Learning Kernels for Processing-In-Memory Architectures
by: Yang, Peiming, et al.
Published: (2025) -
Shared-PIM: Enabling Concurrent Computation and Data Flow for Faster Processing-in-DRAM
by: Mamdouh, Ahmed, et al.
Published: (2024) -
Efficient Sparse Processing-in-Memory Architecture (ESPIM) for Machine Learning Inference
by: He, Mingxuan, et al.
Published: (2024) -
Squire: A General-Purpose Accelerator to Exploit Fine-Grain Parallelism on Dependency-Bound Kernels
by: Langarita, Rubén, et al.
Published: (2025)