Data-aware Dynamic Execution of Irregular Workloads on Heterogeneous Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bai, Zhenyu, Wu, Dan, Dangi, Pranav, Wijerathne, Dhananjaya, Miriyala, Venkata Pavan Kumar, Mitra, Tulika |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HARP: A Taxonomy for Heterogeneous and Hierarchical Processors for Mixed-reuse Workloads
von: Garg, Raveesh, et al.
Veröffentlicht: (2025)
von: Garg, Raveesh, et al.
Veröffentlicht: (2025)
DCRA: A Distributed Chiplet-based Reconfigurable Architecture for Irregular Applications
von: Orenes-Vera, Marcelo, et al.
Veröffentlicht: (2023)
von: Orenes-Vera, Marcelo, et al.
Veröffentlicht: (2023)
Workload-Aware Hardware Accelerator Mining for Distributed Deep Learning Training
von: Adnan, Muhammad, et al.
Veröffentlicht: (2024)
von: Adnan, Muhammad, et al.
Veröffentlicht: (2024)
Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis
von: Chu, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Chu, Xiaoyu, et al.
Veröffentlicht: (2024)
Kitsune: Enabling Dataflow Execution on GPUs
von: Davies, Michael, et al.
Veröffentlicht: (2025)
von: Davies, Michael, et al.
Veröffentlicht: (2025)
Implementation and Evaluation of GBDI Memory Compression Algorithm Using C/C++ on a Broader Range of Workloads
von: Aina, Adeyemi
Veröffentlicht: (2025)
von: Aina, Adeyemi
Veröffentlicht: (2025)
Execution-Centric Characterization of FP8 Matrix Cores, Asynchronous Execution, and Structured Sparsity on AMD MI300A
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)
Achieving Dependability of AI Execution with Radiation Hardened Processors
von: Taquichiri, Carlos Rafael Tordoya, et al.
Veröffentlicht: (2025)
von: Taquichiri, Carlos Rafael Tordoya, et al.
Veröffentlicht: (2025)
XDMA: A Distributed, Extensible DMA Architecture for Layout-Flexible Data Movements in Heterogeneous Multi-Accelerator SoCs
von: Kong, Fanchen, et al.
Veröffentlicht: (2025)
von: Kong, Fanchen, et al.
Veröffentlicht: (2025)
MVDRAM: Enabling GeMV Execution in Unmodified DRAM for Low-Bit LLM Acceleration
von: Kubo, Tatsuya, et al.
Veröffentlicht: (2025)
von: Kubo, Tatsuya, et al.
Veröffentlicht: (2025)
Profiling Concurrent Vision Inference Workloads on NVIDIA Jetson -- Extended
von: Chakraborty, Abhinaba, et al.
Veröffentlicht: (2025)
von: Chakraborty, Abhinaba, et al.
Veröffentlicht: (2025)
Lit Silicon: A Case Where Thermal Imbalance Couples Concurrent Execution in Multiple GPUs
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
Optimizing Offload Performance in Heterogeneous MPSoCs
von: Colagrande, Luca, et al.
Veröffentlicht: (2024)
von: Colagrande, Luca, et al.
Veröffentlicht: (2024)
Next-generation Probabilistic Computing Hardware with 3D MOSAICs, Illusion Scale-up, and Co-design
von: Srimani, Tathagata, et al.
Veröffentlicht: (2024)
von: Srimani, Tathagata, et al.
Veröffentlicht: (2024)
Evaluating Rapid Makespan Predictions for Heterogeneous Systems with Programmable Logic
von: Wilhelm, Martin, et al.
Veröffentlicht: (2025)
von: Wilhelm, Martin, et al.
Veröffentlicht: (2025)
A Heterogeneous Chiplet Architecture for Accelerating End-to-End Transformer Models
von: Sharma, Harsh, et al.
Veröffentlicht: (2023)
von: Sharma, Harsh, et al.
Veröffentlicht: (2023)
EPOCH: Enabling Preemption Operation for Context Saving in Heterogeneous FPGA Systems
von: Malik, Arsalan Ali, et al.
Veröffentlicht: (2025)
von: Malik, Arsalan Ali, et al.
Veröffentlicht: (2025)
Context-aware Simopt-Power: Using structural data with simulation metadata to optimise FPGA designs
von: Wadhwa, Eashan, et al.
Veröffentlicht: (2026)
von: Wadhwa, Eashan, et al.
Veröffentlicht: (2026)
Scheduling Techniques of AI Models on Modern Heterogeneous Edge GPU -- A Critical Review
von: Majeed, Ashiyana Abdul, et al.
Veröffentlicht: (2025)
von: Majeed, Ashiyana Abdul, et al.
Veröffentlicht: (2025)
Proteus: Enabling High-Performance Processing-Using-DRAM with Dynamic Bit-Precision, Adaptive Data Representation, and Flexible Arithmetic
von: Oliveira, Geraldo F., et al.
Veröffentlicht: (2025)
von: Oliveira, Geraldo F., et al.
Veröffentlicht: (2025)
HgPCN: A Heterogeneous Architecture for E2E Embedded Point Cloud Inference
von: Gao, Yiming, et al.
Veröffentlicht: (2025)
von: Gao, Yiming, et al.
Veröffentlicht: (2025)
GreenLLM: Disaggregating Large Language Model Serving on Heterogeneous GPUs for Lower Carbon Emissions
von: Shi, Tianyao, et al.
Veröffentlicht: (2024)
von: Shi, Tianyao, et al.
Veröffentlicht: (2024)
A Survey of Real-time Scheduling on Accelerator-based Heterogeneous Architecture for Time Critical Applications
von: Zou, An, et al.
Veröffentlicht: (2025)
von: Zou, An, et al.
Veröffentlicht: (2025)
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
von: Xu, Weihong, et al.
Veröffentlicht: (2025)
von: Xu, Weihong, et al.
Veröffentlicht: (2025)
A Reliable, Time-Predictable Heterogeneous SoC for AI-Enhanced Mixed-Criticality Edge Applications
von: Garofalo, Angelo, et al.
Veröffentlicht: (2025)
von: Garofalo, Angelo, et al.
Veröffentlicht: (2025)
Compiler Support for Speculation in Decoupled Access/Execute Architectures
von: Szafarczyk, Robert, et al.
Veröffentlicht: (2025)
von: Szafarczyk, Robert, et al.
Veröffentlicht: (2025)
Dynamic Simultaneous Multithreaded Architecture
von: Ortiz-Arroyo, Daniel, et al.
Veröffentlicht: (2024)
von: Ortiz-Arroyo, Daniel, et al.
Veröffentlicht: (2024)
Revisiting Computational Storage for Data Integrity and Security
von: Shi, Chao, et al.
Veröffentlicht: (2025)
von: Shi, Chao, et al.
Veröffentlicht: (2025)
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
von: Zhang, Qijun, et al.
Veröffentlicht: (2026)
von: Zhang, Qijun, et al.
Veröffentlicht: (2026)
Balanced Data Placement for GEMV Acceleration with Processing-In-Memory
von: Ibrahim, Mohamed Assem, et al.
Veröffentlicht: (2024)
von: Ibrahim, Mohamed Assem, et al.
Veröffentlicht: (2024)
Accelerating Data Chunking in Deduplication Systems using Vector Instructions
von: Udayashankar, Sreeharsha, et al.
Veröffentlicht: (2025)
von: Udayashankar, Sreeharsha, et al.
Veröffentlicht: (2025)
DP-HLS: A High-Level Synthesis Framework for Accelerating Dynamic Programming Algorithms in Bioinformatics
von: Cao, Yingqi, et al.
Veröffentlicht: (2024)
von: Cao, Yingqi, et al.
Veröffentlicht: (2024)
Torrent: A Distributed DMA for Efficient and Flexible Point-to-Multipoint Data Movement
von: Deng, Yunhao, et al.
Veröffentlicht: (2025)
von: Deng, Yunhao, et al.
Veröffentlicht: (2025)
RAPID-Graph: Recursive All-Pairs Shortest Paths Using Processing-in-Memory for Dynamic Programming on Graphs
von: Chen, Yanru, et al.
Veröffentlicht: (2025)
von: Chen, Yanru, et al.
Veröffentlicht: (2025)
FpgaHub: Fpga-centric Hyper-heterogeneous Computing Platform for Big Data Analytics
von: Wang, Zeke, et al.
Veröffentlicht: (2025)
von: Wang, Zeke, et al.
Veröffentlicht: (2025)
EDEA: Efficient Dual-Engine Accelerator for Depthwise Separable Convolution with Direct Data Transfer
von: Chen, Yi, et al.
Veröffentlicht: (2025)
von: Chen, Yi, et al.
Veröffentlicht: (2025)
Conduit: Programmer-Transparent Near-Data Processing Using Multiple Compute-Capable Resources in Solid State Drives
von: Nadig, Rakesh, et al.
Veröffentlicht: (2026)
von: Nadig, Rakesh, et al.
Veröffentlicht: (2026)
Performance Modeling and Workload Analysis of Distributed Large Language Model Training and Inference
von: Kundu, Joyjit, et al.
Veröffentlicht: (2024)
von: Kundu, Joyjit, et al.
Veröffentlicht: (2024)
MIMDRAM: An End-to-End Processing-Using-DRAM System for High-Throughput, Energy-Efficient and Programmer-Transparent Multiple-Instruction Multiple-Data Processing
von: Oliveira, Geraldo F., et al.
Veröffentlicht: (2024)
von: Oliveira, Geraldo F., et al.
Veröffentlicht: (2024)
TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
von: Kwak, Hyunseok, et al.
Veröffentlicht: (2025)
von: Kwak, Hyunseok, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
HARP: A Taxonomy for Heterogeneous and Hierarchical Processors for Mixed-reuse Workloads
von: Garg, Raveesh, et al.
Veröffentlicht: (2025) -
DCRA: A Distributed Chiplet-based Reconfigurable Architecture for Irregular Applications
von: Orenes-Vera, Marcelo, et al.
Veröffentlicht: (2023) -
Workload-Aware Hardware Accelerator Mining for Distributed Deep Learning Training
von: Adnan, Muhammad, et al.
Veröffentlicht: (2024) -
Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis
von: Chu, Xiaoyu, et al.
Veröffentlicht: (2024) -
Kitsune: Enabling Dataflow Execution on GPUs
von: Davies, Michael, et al.
Veröffentlicht: (2025)