Optimizing Offload Performance in Heterogeneous MPSoCs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Colagrande, Luca, Benini, Luca |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Taming Offload Overheads in a Massively Parallel Open-Source RISC-V MPSoC: Analysis and Optimization
par: Colagrande, Luca, et autres
Publié: (2025)
par: Colagrande, Luca, et autres
Publié: (2025)
A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators
par: Colagrande, Luca, et autres
Publié: (2026)
par: Colagrande, Luca, et autres
Publié: (2026)
An End-to-End DNN Inference Framework for the SpiNNaker2 Neuromorphic MPSoC
par: Jobst, Matthias, et autres
Publié: (2025)
par: Jobst, Matthias, et autres
Publié: (2025)
TCDM Burst Access: Breaking the Bandwidth Barrier in Shared-L1 RVV Clusters Beyond 1000 FPUs
par: Shen, Diyou, et autres
Publié: (2025)
par: Shen, Diyou, et autres
Publié: (2025)
MemPool Flavors: Between Versatility and Specialization in a RISC-V Manycore Cluster
par: Mazzola, Sergio, et autres
Publié: (2025)
par: Mazzola, Sergio, et autres
Publié: (2025)
MATCHA: Efficient Deployment of Deep Neural Networks on Multi-Accelerator Heterogeneous Edge SoCs
par: Russo, Enrico, et autres
Publié: (2026)
par: Russo, Enrico, et autres
Publié: (2026)
Design in Tiles: Automating GEMM Deployment on Tile-Based Many-PE Accelerators
par: Shen, Aofeng, et autres
Publié: (2025)
par: Shen, Aofeng, et autres
Publié: (2025)
A Reliable, Time-Predictable Heterogeneous SoC for AI-Enhanced Mixed-Criticality Edge Applications
par: Garofalo, Angelo, et autres
Publié: (2025)
par: Garofalo, Angelo, et autres
Publié: (2025)
TeraPool-SDR: An 1.89TOPS 1024 RV-Cores 4MiB Shared-L1 Cluster for Next-Generation Open-Source Software-Defined Radios
par: Zhang, Yichao, et autres
Publié: (2024)
par: Zhang, Yichao, et autres
Publié: (2024)
TeraPool: A Physical Design Aware, 1024 RISC-V Cores Shared-L1-Memory Scaled-up Cluster Design with High Bandwidth Main Memory Link
par: Zhang, Yichao, et autres
Publié: (2026)
par: Zhang, Yichao, et autres
Publié: (2026)
XDMA: A Distributed, Extensible DMA Architecture for Layout-Flexible Data Movements in Heterogeneous Multi-Accelerator SoCs
par: Kong, Fanchen, et autres
Publié: (2025)
par: Kong, Fanchen, et autres
Publié: (2025)
Understanding Bottlenecks for Efficiently Serving LLM Inference With KV Offloading
par: Meng, William, et autres
Publié: (2025)
par: Meng, William, et autres
Publié: (2025)
Knowledge-Guided Attention-Inspired Learning for Task Offloading in Vehicle Edge Computing
par: Ma, Ke, et autres
Publié: (2025)
par: Ma, Ke, et autres
Publié: (2025)
DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication
par: Pati, Suchita, et autres
Publié: (2025)
par: Pati, Suchita, et autres
Publié: (2025)
Efficient deadlock avoidance for 2D mesh NoCs that use OQ or VOQ routers
par: Papaphilippou, Philippos, et autres
Publié: (2023)
par: Papaphilippou, Philippos, et autres
Publié: (2023)
Optimizing Foundation Model Inference on a Many-tiny-core Open-source RISC-V Platform
par: Potocnik, Viviane, et autres
Publié: (2024)
par: Potocnik, Viviane, et autres
Publié: (2024)
Evaluating Rapid Makespan Predictions for Heterogeneous Systems with Programmable Logic
par: Wilhelm, Martin, et autres
Publié: (2025)
par: Wilhelm, Martin, et autres
Publié: (2025)
Data-aware Dynamic Execution of Irregular Workloads on Heterogeneous Systems
par: Bai, Zhenyu, et autres
Publié: (2025)
par: Bai, Zhenyu, et autres
Publié: (2025)
Enabling Time-Aware Priority Traffic Management over Distributed FPGA Nodes
par: Scionti, Alberto, et autres
Publié: (2025)
par: Scionti, Alberto, et autres
Publié: (2025)
A Heterogeneous Chiplet Architecture for Accelerating End-to-End Transformer Models
par: Sharma, Harsh, et autres
Publié: (2023)
par: Sharma, Harsh, et autres
Publié: (2023)
EPOCH: Enabling Preemption Operation for Context Saving in Heterogeneous FPGA Systems
par: Malik, Arsalan Ali, et autres
Publié: (2025)
par: Malik, Arsalan Ali, et autres
Publié: (2025)
HARP: A Taxonomy for Heterogeneous and Hierarchical Processors for Mixed-reuse Workloads
par: Garg, Raveesh, et autres
Publié: (2025)
par: Garg, Raveesh, et autres
Publié: (2025)
EPAC: The Last Dance
par: Mantovani, Filippo, et autres
Publié: (2026)
par: Mantovani, Filippo, et autres
Publié: (2026)
Scheduling Techniques of AI Models on Modern Heterogeneous Edge GPU -- A Critical Review
par: Majeed, Ashiyana Abdul, et autres
Publié: (2025)
par: Majeed, Ashiyana Abdul, et autres
Publié: (2025)
GreenLLM: Disaggregating Large Language Model Serving on Heterogeneous GPUs for Lower Carbon Emissions
par: Shi, Tianyao, et autres
Publié: (2024)
par: Shi, Tianyao, et autres
Publié: (2024)
HgPCN: A Heterogeneous Architecture for E2E Embedded Point Cloud Inference
par: Gao, Yiming, et autres
Publié: (2025)
par: Gao, Yiming, et autres
Publié: (2025)
A Survey of Real-time Scheduling on Accelerator-based Heterogeneous Architecture for Time Critical Applications
par: Zou, An, et autres
Publié: (2025)
par: Zou, An, et autres
Publié: (2025)
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
par: Xu, Weihong, et autres
Publié: (2025)
par: Xu, Weihong, et autres
Publié: (2025)
Optimizing Task Scheduling in Fog Computing with Deadline Awareness
par: Sirjani, Mohammad Sadegh, et autres
Publié: (2025)
par: Sirjani, Mohammad Sadegh, et autres
Publié: (2025)
iHAC: A Hybrid Cluster Architecture for Enhanced Performance and Resilience
par: Muntaka, Siddique Abubakr, et autres
Publié: (2026)
par: Muntaka, Siddique Abubakr, et autres
Publié: (2026)
Microbenchmark-Driven Analytical Performance Modeling Across Modern GPU Architectures
par: Jarmusch, Aaron, et autres
Publié: (2026)
par: Jarmusch, Aaron, et autres
Publié: (2026)
NetSmith: An Optimization Framework for Machine-Discovered Network Topologies
par: Green, Conor, et autres
Publié: (2024)
par: Green, Conor, et autres
Publié: (2024)
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
par: Agrawal, Anirudha, et autres
Publié: (2024)
par: Agrawal, Anirudha, et autres
Publié: (2024)
Optimizing Distributed ML Communication with Fused Computation-Collective Operations
par: Punniyamurthy, Kishore, et autres
Publié: (2023)
par: Punniyamurthy, Kishore, et autres
Publié: (2023)
MOFCO: Mobility- and Migration-Aware Task Offloading in Three-Layer Fog Computing Environments
par: Mahdizadeh, Soheil, et autres
Publié: (2025)
par: Mahdizadeh, Soheil, et autres
Publié: (2025)
Multi-Partner Project: Multi-GPU Performance Portability Analysis for CFD Simulations at Scale
par: Eleftherakis, Panagiotis-Eleftherios, et autres
Publié: (2026)
par: Eleftherakis, Panagiotis-Eleftherios, et autres
Publié: (2026)
CMDS: Cross-layer Dataflow Optimization for DNN Accelerators Exploiting Multi-bank Memories
par: Shi, Man, et autres
Publié: (2024)
par: Shi, Man, et autres
Publié: (2024)
Optimizing Communication for Latency Sensitive HPC Applications on up to 48 FPGAs Using ACCL
par: Meyer, Marius, et autres
Publié: (2024)
par: Meyer, Marius, et autres
Publié: (2024)
PULSAR: Simultaneous Many-Row Activation for Reliable and High-Performance Computing in Off-the-Shelf DRAM Chips
par: Yuksel, Ismail Emir, et autres
Publié: (2023)
par: Yuksel, Ismail Emir, et autres
Publié: (2023)
Proteus: Enabling High-Performance Processing-Using-DRAM with Dynamic Bit-Precision, Adaptive Data Representation, and Flexible Arithmetic
par: Oliveira, Geraldo F., et autres
Publié: (2025)
par: Oliveira, Geraldo F., et autres
Publié: (2025)
Documents similaires
-
Taming Offload Overheads in a Massively Parallel Open-Source RISC-V MPSoC: Analysis and Optimization
par: Colagrande, Luca, et autres
Publié: (2025) -
A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators
par: Colagrande, Luca, et autres
Publié: (2026) -
An End-to-End DNN Inference Framework for the SpiNNaker2 Neuromorphic MPSoC
par: Jobst, Matthias, et autres
Publié: (2025) -
TCDM Burst Access: Breaking the Bandwidth Barrier in Shared-L1 RVV Clusters Beyond 1000 FPUs
par: Shen, Diyou, et autres
Publié: (2025) -
MemPool Flavors: Between Versatility and Specialization in a RISC-V Manycore Cluster
par: Mazzola, Sergio, et autres
Publié: (2025)