Taming Offload Overheads in a Massively Parallel Open-Source RISC-V MPSoC: Analysis and Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Colagrande, Luca, Benini, Luca |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Optimizing Offload Performance in Heterogeneous MPSoCs
von: Colagrande, Luca, et al.
Veröffentlicht: (2024)
von: Colagrande, Luca, et al.
Veröffentlicht: (2024)
A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators
von: Colagrande, Luca, et al.
Veröffentlicht: (2026)
von: Colagrande, Luca, et al.
Veröffentlicht: (2026)
MemPool Flavors: Between Versatility and Specialization in a RISC-V Manycore Cluster
von: Mazzola, Sergio, et al.
Veröffentlicht: (2025)
von: Mazzola, Sergio, et al.
Veröffentlicht: (2025)
An End-to-End DNN Inference Framework for the SpiNNaker2 Neuromorphic MPSoC
von: Jobst, Matthias, et al.
Veröffentlicht: (2025)
von: Jobst, Matthias, et al.
Veröffentlicht: (2025)
Optimizing Foundation Model Inference on a Many-tiny-core Open-source RISC-V Platform
von: Potocnik, Viviane, et al.
Veröffentlicht: (2024)
von: Potocnik, Viviane, et al.
Veröffentlicht: (2024)
TeraPool: A Physical Design Aware, 1024 RISC-V Cores Shared-L1-Memory Scaled-up Cluster Design with High Bandwidth Main Memory Link
von: Zhang, Yichao, et al.
Veröffentlicht: (2026)
von: Zhang, Yichao, et al.
Veröffentlicht: (2026)
TeraPool-SDR: An 1.89TOPS 1024 RV-Cores 4MiB Shared-L1 Cluster for Next-Generation Open-Source Software-Defined Radios
von: Zhang, Yichao, et al.
Veröffentlicht: (2024)
von: Zhang, Yichao, et al.
Veröffentlicht: (2024)
TCDM Burst Access: Breaking the Bandwidth Barrier in Shared-L1 RVV Clusters Beyond 1000 FPUs
von: Shen, Diyou, et al.
Veröffentlicht: (2025)
von: Shen, Diyou, et al.
Veröffentlicht: (2025)
Efficient Architecture for RISC-V Vector Memory Access
von: Guan, Hongyi, et al.
Veröffentlicht: (2025)
von: Guan, Hongyi, et al.
Veröffentlicht: (2025)
RISC-V Word-Size Modular Instructions for Residue Number Systems
von: Didier, Laurent-Stéphane, et al.
Veröffentlicht: (2024)
von: Didier, Laurent-Stéphane, et al.
Veröffentlicht: (2024)
Design in Tiles: Automating GEMM Deployment on Tile-Based Many-PE Accelerators
von: Shen, Aofeng, et al.
Veröffentlicht: (2025)
von: Shen, Aofeng, et al.
Veröffentlicht: (2025)
A Reliable, Time-Predictable Heterogeneous SoC for AI-Enhanced Mixed-Criticality Edge Applications
von: Garofalo, Angelo, et al.
Veröffentlicht: (2025)
von: Garofalo, Angelo, et al.
Veröffentlicht: (2025)
Open Challenges for a Production-ready Cloud Environment on top of RISC-V hardware
von: Call, Aaron, et al.
Veröffentlicht: (2025)
von: Call, Aaron, et al.
Veröffentlicht: (2025)
Analyzing Reverse Address Translation Overheads in Multi-GPU Scale-Up Pods
von: Fatima, Amel, et al.
Veröffentlicht: (2026)
von: Fatima, Amel, et al.
Veröffentlicht: (2026)
Understanding Bottlenecks for Efficiently Serving LLM Inference With KV Offloading
von: Meng, William, et al.
Veröffentlicht: (2025)
von: Meng, William, et al.
Veröffentlicht: (2025)
Switchboard: An Open-Source Framework for Modular Simulation of Large Hardware Systems
von: Herbst, Steven, et al.
Veröffentlicht: (2024)
von: Herbst, Steven, et al.
Veröffentlicht: (2024)
Knowledge-Guided Attention-Inspired Learning for Task Offloading in Vehicle Edge Computing
von: Ma, Ke, et al.
Veröffentlicht: (2025)
von: Ma, Ke, et al.
Veröffentlicht: (2025)
DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication
von: Pati, Suchita, et al.
Veröffentlicht: (2025)
von: Pati, Suchita, et al.
Veröffentlicht: (2025)
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
von: Zhou, Zhuoshan, et al.
Veröffentlicht: (2026)
von: Zhou, Zhuoshan, et al.
Veröffentlicht: (2026)
Parendi: Thousand-Way Parallel RTL Simulation
von: Emami, Mahyar, et al.
Veröffentlicht: (2024)
von: Emami, Mahyar, et al.
Veröffentlicht: (2024)
Datapath Combinational Equivalence Checking With Hybrid Sweeping Engines and Parallelization
von: Chen, Zhihan, et al.
Veröffentlicht: (2024)
von: Chen, Zhihan, et al.
Veröffentlicht: (2024)
How Fast Can Graph Computations Go on Fine-grained Parallel Architectures
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
Deep Learning and Machine Learning with GPGPU and CUDA: Unlocking the Power of Parallel Computing
von: Li, Ming, et al.
Veröffentlicht: (2024)
von: Li, Ming, et al.
Veröffentlicht: (2024)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
von: Zhang, Chen, et al.
Veröffentlicht: (2026)
von: Zhang, Chen, et al.
Veröffentlicht: (2026)
Enabling Time-Aware Priority Traffic Management over Distributed FPGA Nodes
von: Scionti, Alberto, et al.
Veröffentlicht: (2025)
von: Scionti, Alberto, et al.
Veröffentlicht: (2025)
EPAC: The Last Dance
von: Mantovani, Filippo, et al.
Veröffentlicht: (2026)
von: Mantovani, Filippo, et al.
Veröffentlicht: (2026)
Parallelizing a modern GPU simulator
von: Huerta, Rodrigo, et al.
Veröffentlicht: (2025)
von: Huerta, Rodrigo, et al.
Veröffentlicht: (2025)
RapidOMS: FPGA-based Open Modification Spectral Library Searching with HD Computing
von: Pinge, Sumukh, et al.
Veröffentlicht: (2024)
von: Pinge, Sumukh, et al.
Veröffentlicht: (2024)
Implementation and Evaluation of GBDI Memory Compression Algorithm Using C/C++ on a Broader Range of Workloads
von: Aina, Adeyemi
Veröffentlicht: (2025)
von: Aina, Adeyemi
Veröffentlicht: (2025)
GigaAPI for GPU Parallelization
von: Suvarna, M., et al.
Veröffentlicht: (2025)
von: Suvarna, M., et al.
Veröffentlicht: (2025)
FinGraV: Methodology for Fine-Grain GPU Power Visibility and Insights
von: Singhania, Varsha, et al.
Veröffentlicht: (2024)
von: Singhania, Varsha, et al.
Veröffentlicht: (2024)
Optimizing Task Scheduling in Fog Computing with Deadline Awareness
von: Sirjani, Mohammad Sadegh, et al.
Veröffentlicht: (2025)
von: Sirjani, Mohammad Sadegh, et al.
Veröffentlicht: (2025)
NetSmith: An Optimization Framework for Machine-Discovered Network Topologies
von: Green, Conor, et al.
Veröffentlicht: (2024)
von: Green, Conor, et al.
Veröffentlicht: (2024)
Optimizing Distributed ML Communication with Fused Computation-Collective Operations
von: Punniyamurthy, Kishore, et al.
Veröffentlicht: (2023)
von: Punniyamurthy, Kishore, et al.
Veröffentlicht: (2023)
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
von: Agrawal, Anirudha, et al.
Veröffentlicht: (2024)
von: Agrawal, Anirudha, et al.
Veröffentlicht: (2024)
MOFCO: Mobility- and Migration-Aware Task Offloading in Three-Layer Fog Computing Environments
von: Mahdizadeh, Soheil, et al.
Veröffentlicht: (2025)
von: Mahdizadeh, Soheil, et al.
Veröffentlicht: (2025)
MATCHA: Efficient Deployment of Deep Neural Networks on Multi-Accelerator Heterogeneous Edge SoCs
von: Russo, Enrico, et al.
Veröffentlicht: (2026)
von: Russo, Enrico, et al.
Veröffentlicht: (2026)
Cloud-Native Operation of Roadside Infrastructure Enabling Demand-Driven Collective Perception via V2X
von: Zanger, Lukas, et al.
Veröffentlicht: (2026)
von: Zanger, Lukas, et al.
Veröffentlicht: (2026)
CMDS: Cross-layer Dataflow Optimization for DNN Accelerators Exploiting Multi-bank Memories
von: Shi, Man, et al.
Veröffentlicht: (2024)
von: Shi, Man, et al.
Veröffentlicht: (2024)
Optimizing Communication for Latency Sensitive HPC Applications on up to 48 FPGAs Using ACCL
von: Meyer, Marius, et al.
Veröffentlicht: (2024)
von: Meyer, Marius, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Optimizing Offload Performance in Heterogeneous MPSoCs
von: Colagrande, Luca, et al.
Veröffentlicht: (2024) -
A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators
von: Colagrande, Luca, et al.
Veröffentlicht: (2026) -
MemPool Flavors: Between Versatility and Specialization in a RISC-V Manycore Cluster
von: Mazzola, Sergio, et al.
Veröffentlicht: (2025) -
An End-to-End DNN Inference Framework for the SpiNNaker2 Neuromorphic MPSoC
von: Jobst, Matthias, et al.
Veröffentlicht: (2025) -
Optimizing Foundation Model Inference on a Many-tiny-core Open-source RISC-V Platform
von: Potocnik, Viviane, et al.
Veröffentlicht: (2024)