A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Colagrande, Luca, Leone, Lorenzo, Wu, Chen, Fischer, Tim, Roth, Raphael, Benini, Luca |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Taming Offload Overheads in a Massively Parallel Open-Source RISC-V MPSoC: Analysis and Optimization
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
Optimizing Offload Performance in Heterogeneous MPSoCs
von: Colagrande, Luca, et al.
Veröffentlicht: (2024)
von: Colagrande, Luca, et al.
Veröffentlicht: (2024)
Design in Tiles: Automating GEMM Deployment on Tile-Based Many-PE Accelerators
von: Shen, Aofeng, et al.
Veröffentlicht: (2025)
von: Shen, Aofeng, et al.
Veröffentlicht: (2025)
TCDM Burst Access: Breaking the Bandwidth Barrier in Shared-L1 RVV Clusters Beyond 1000 FPUs
von: Shen, Diyou, et al.
Veröffentlicht: (2025)
von: Shen, Diyou, et al.
Veröffentlicht: (2025)
TeraPool: A Physical Design Aware, 1024 RISC-V Cores Shared-L1-Memory Scaled-up Cluster Design with High Bandwidth Main Memory Link
von: Zhang, Yichao, et al.
Veröffentlicht: (2026)
von: Zhang, Yichao, et al.
Veröffentlicht: (2026)
MemPool Flavors: Between Versatility and Specialization in a RISC-V Manycore Cluster
von: Mazzola, Sergio, et al.
Veröffentlicht: (2025)
von: Mazzola, Sergio, et al.
Veröffentlicht: (2025)
Efficient deadlock avoidance for 2D mesh NoCs that use OQ or VOQ routers
von: Papaphilippou, Philippos, et al.
Veröffentlicht: (2023)
von: Papaphilippou, Philippos, et al.
Veröffentlicht: (2023)
TeraPool-SDR: An 1.89TOPS 1024 RV-Cores 4MiB Shared-L1 Cluster for Next-Generation Open-Source Software-Defined Radios
von: Zhang, Yichao, et al.
Veröffentlicht: (2024)
von: Zhang, Yichao, et al.
Veröffentlicht: (2024)
Optimizing Distributed ML Communication with Fused Computation-Collective Operations
von: Punniyamurthy, Kishore, et al.
Veröffentlicht: (2023)
von: Punniyamurthy, Kishore, et al.
Veröffentlicht: (2023)
A Reliable, Time-Predictable Heterogeneous SoC for AI-Enhanced Mixed-Criticality Edge Applications
von: Garofalo, Angelo, et al.
Veröffentlicht: (2025)
von: Garofalo, Angelo, et al.
Veröffentlicht: (2025)
Optimizing Foundation Model Inference on a Many-tiny-core Open-source RISC-V Platform
von: Potocnik, Viviane, et al.
Veröffentlicht: (2024)
von: Potocnik, Viviane, et al.
Veröffentlicht: (2024)
Leveraging SIMD for Accelerating Large-number Arithmetic
von: Das, Subhrajit, et al.
Veröffentlicht: (2026)
von: Das, Subhrajit, et al.
Veröffentlicht: (2026)
Accelerating Triangle Counting with Real Processing-in-Memory Systems
von: Asquini, Lorenzo, et al.
Veröffentlicht: (2025)
von: Asquini, Lorenzo, et al.
Veröffentlicht: (2025)
MIMDRAM: An End-to-End Processing-Using-DRAM System for High-Throughput, Energy-Efficient and Programmer-Transparent Multiple-Instruction Multiple-Data Processing
von: Oliveira, Geraldo F., et al.
Veröffentlicht: (2024)
von: Oliveira, Geraldo F., et al.
Veröffentlicht: (2024)
MATCHA: Efficient Deployment of Deep Neural Networks on Multi-Accelerator Heterogeneous Edge SoCs
von: Russo, Enrico, et al.
Veröffentlicht: (2026)
von: Russo, Enrico, et al.
Veröffentlicht: (2026)
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
von: Agrawal, Anirudha, et al.
Veröffentlicht: (2024)
von: Agrawal, Anirudha, et al.
Veröffentlicht: (2024)
SAGe: A Lightweight Algorithm-Architecture Co-Design for Mitigating the Data Preparation Bottleneck in Large-Scale Genome Sequence Analysis
von: Ghiasi, Nika Mansouri, et al.
Veröffentlicht: (2025)
von: Ghiasi, Nika Mansouri, et al.
Veröffentlicht: (2025)
EPAC: The Last Dance
von: Mantovani, Filippo, et al.
Veröffentlicht: (2026)
von: Mantovani, Filippo, et al.
Veröffentlicht: (2026)
DP-HLS: A High-Level Synthesis Framework for Accelerating Dynamic Programming Algorithms in Bioinformatics
von: Cao, Yingqi, et al.
Veröffentlicht: (2024)
von: Cao, Yingqi, et al.
Veröffentlicht: (2024)
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
von: Xu, Weihong, et al.
Veröffentlicht: (2025)
von: Xu, Weihong, et al.
Veröffentlicht: (2025)
DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication
von: Pati, Suchita, et al.
Veröffentlicht: (2025)
von: Pati, Suchita, et al.
Veröffentlicht: (2025)
LFOC: A Lightweight Fairness-Oriented Cache Clustering Policy for Commodity Multicores
von: García-García, Adrián, et al.
Veröffentlicht: (2024)
von: García-García, Adrián, et al.
Veröffentlicht: (2024)
Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis
von: Chu, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Chu, Xiaoyu, et al.
Veröffentlicht: (2024)
SpeedMalloc: Improving Multi-threaded Applications via a Lightweight Core for Memory Allocation
von: Li, Ruihao, et al.
Veröffentlicht: (2025)
von: Li, Ruihao, et al.
Veröffentlicht: (2025)
Sparse MTTKRP Acceleration for Tensor Decomposition on GPU
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2024)
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2024)
Balanced Data Placement for GEMV Acceleration with Processing-In-Memory
von: Ibrahim, Mohamed Assem, et al.
Veröffentlicht: (2024)
von: Ibrahim, Mohamed Assem, et al.
Veröffentlicht: (2024)
Enabling Time-Aware Priority Traffic Management over Distributed FPGA Nodes
von: Scionti, Alberto, et al.
Veröffentlicht: (2025)
von: Scionti, Alberto, et al.
Veröffentlicht: (2025)
Conduit: Programmer-Transparent Near-Data Processing Using Multiple Compute-Capable Resources in Solid State Drives
von: Nadig, Rakesh, et al.
Veröffentlicht: (2026)
von: Nadig, Rakesh, et al.
Veröffentlicht: (2026)
Accelerating Data Chunking in Deduplication Systems using Vector Instructions
von: Udayashankar, Sreeharsha, et al.
Veröffentlicht: (2025)
von: Udayashankar, Sreeharsha, et al.
Veröffentlicht: (2025)
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
von: Zhang, Qijun, et al.
Veröffentlicht: (2026)
von: Zhang, Qijun, et al.
Veröffentlicht: (2026)
COMET: A Framework for Modeling Compound Operation Dataflows with Explicit Collectives
von: Negi, Shubham, et al.
Veröffentlicht: (2025)
von: Negi, Shubham, et al.
Veröffentlicht: (2025)
A Heterogeneous Chiplet Architecture for Accelerating End-to-End Transformer Models
von: Sharma, Harsh, et al.
Veröffentlicht: (2023)
von: Sharma, Harsh, et al.
Veröffentlicht: (2023)
Application Experiences on a GPU-Accelerated Arm-based HPC Testbed
von: Elwasif, Wael, et al.
Veröffentlicht: (2022)
von: Elwasif, Wael, et al.
Veröffentlicht: (2022)
Workload-Aware Hardware Accelerator Mining for Distributed Deep Learning Training
von: Adnan, Muhammad, et al.
Veröffentlicht: (2024)
von: Adnan, Muhammad, et al.
Veröffentlicht: (2024)
TAPA-CS: Enabling Scalable Accelerator Design on Distributed HBM-FPGAs
von: Prakriya, Neha, et al.
Veröffentlicht: (2023)
von: Prakriya, Neha, et al.
Veröffentlicht: (2023)
The DEEP-ER project: I/O and resiliency extensions for the Cluster-Booster architecture
von: Kreuzer, Anke, et al.
Veröffentlicht: (2019)
von: Kreuzer, Anke, et al.
Veröffentlicht: (2019)
FLEX: Leveraging FPGA-CPU Synergy for Mixed-Cell-Height Legalization Acceleration
von: Liu, Xingyu, et al.
Veröffentlicht: (2025)
von: Liu, Xingyu, et al.
Veröffentlicht: (2025)
An Evaluation and Comparison of GPU Hardware and Solver Libraries for Accelerating the OPM Flow Reservoir Simulator
von: Qiu, Tong Dong, et al.
Veröffentlicht: (2023)
von: Qiu, Tong Dong, et al.
Veröffentlicht: (2023)
MVDRAM: Enabling GeMV Execution in Unmodified DRAM for Low-Bit LLM Acceleration
von: Kubo, Tatsuya, et al.
Veröffentlicht: (2025)
von: Kubo, Tatsuya, et al.
Veröffentlicht: (2025)
CMDS: Cross-layer Dataflow Optimization for DNN Accelerators Exploiting Multi-bank Memories
von: Shi, Man, et al.
Veröffentlicht: (2024)
von: Shi, Man, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Taming Offload Overheads in a Massively Parallel Open-Source RISC-V MPSoC: Analysis and Optimization
von: Colagrande, Luca, et al.
Veröffentlicht: (2025) -
Optimizing Offload Performance in Heterogeneous MPSoCs
von: Colagrande, Luca, et al.
Veröffentlicht: (2024) -
Design in Tiles: Automating GEMM Deployment on Tile-Based Many-PE Accelerators
von: Shen, Aofeng, et al.
Veröffentlicht: (2025) -
TCDM Burst Access: Breaking the Bandwidth Barrier in Shared-L1 RVV Clusters Beyond 1000 FPUs
von: Shen, Diyou, et al.
Veröffentlicht: (2025) -
TeraPool: A Physical Design Aware, 1024 RISC-V Cores Shared-L1-Memory Scaled-up Cluster Design with High Bandwidth Main Memory Link
von: Zhang, Yichao, et al.
Veröffentlicht: (2026)