Accelerating Precise End-to-End Simulation: Latency-Sensitive Many-core System Modeling
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Yinrong, Fu, Zexin, Zhang, Yichao, Haugou, Germain, Zhang, Chi, Bertuletti, Marco, Wang, Bowen, Benini, Luca |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
TCDM Burst Access: Breaking the Bandwidth Barrier in Shared-L1 RVV Clusters Beyond 1000 FPUs
di: Shen, Diyou, et al.
Pubblicazione: (2025)
di: Shen, Diyou, et al.
Pubblicazione: (2025)
Optimizing Foundation Model Inference on a Many-tiny-core Open-source RISC-V Platform
di: Potocnik, Viviane, et al.
Pubblicazione: (2024)
di: Potocnik, Viviane, et al.
Pubblicazione: (2024)
MemPool Flavors: Between Versatility and Specialization in a RISC-V Manycore Cluster
di: Mazzola, Sergio, et al.
Pubblicazione: (2025)
di: Mazzola, Sergio, et al.
Pubblicazione: (2025)
TeraPool: A Physical Design Aware, 1024 RISC-V Cores Shared-L1-Memory Scaled-up Cluster Design with High Bandwidth Main Memory Link
di: Zhang, Yichao, et al.
Pubblicazione: (2026)
di: Zhang, Yichao, et al.
Pubblicazione: (2026)
Lincoln AI Computing Survey (LAICS) and Trends
di: Reuther, Albert, et al.
Pubblicazione: (2025)
di: Reuther, Albert, et al.
Pubblicazione: (2025)
TeraPool-SDR: An 1.89TOPS 1024 RV-Cores 4MiB Shared-L1 Cluster for Next-Generation Open-Source Software-Defined Radios
di: Zhang, Yichao, et al.
Pubblicazione: (2024)
di: Zhang, Yichao, et al.
Pubblicazione: (2024)
Design in Tiles: Automating GEMM Deployment on Tile-Based Many-PE Accelerators
di: Shen, Aofeng, et al.
Pubblicazione: (2025)
di: Shen, Aofeng, et al.
Pubblicazione: (2025)
Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs
di: Peng, Hongwu, et al.
Pubblicazione: (2023)
di: Peng, Hongwu, et al.
Pubblicazione: (2023)
Toward a Universal GPU Instruction Set Architecture: A Cross-Vendor Analysis of Hardware-Invariant Computational Primitives in Parallel Processors
di: Abraham, Ojima, et al.
Pubblicazione: (2026)
di: Abraham, Ojima, et al.
Pubblicazione: (2026)
Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference Environments
di: Iliakopoulou, Nikoleta, et al.
Pubblicazione: (2024)
di: Iliakopoulou, Nikoleta, et al.
Pubblicazione: (2024)
A Heterogeneous Chiplet Architecture for Accelerating End-to-End Transformer Models
di: Sharma, Harsh, et al.
Pubblicazione: (2023)
di: Sharma, Harsh, et al.
Pubblicazione: (2023)
Optimizing Offload Performance in Heterogeneous MPSoCs
di: Colagrande, Luca, et al.
Pubblicazione: (2024)
di: Colagrande, Luca, et al.
Pubblicazione: (2024)
Taming Offload Overheads in a Massively Parallel Open-Source RISC-V MPSoC: Analysis and Optimization
di: Colagrande, Luca, et al.
Pubblicazione: (2025)
di: Colagrande, Luca, et al.
Pubblicazione: (2025)
T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives
di: Pati, Suchita, et al.
Pubblicazione: (2024)
di: Pati, Suchita, et al.
Pubblicazione: (2024)
Mitigating the Memory Bottleneck with Machine Learning-Driven and Data-Aware Microarchitectural Techniques
di: Bera, Rahul
Pubblicazione: (2026)
di: Bera, Rahul
Pubblicazione: (2026)
A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators
di: Colagrande, Luca, et al.
Pubblicazione: (2026)
di: Colagrande, Luca, et al.
Pubblicazione: (2026)
LLaMCAT: Optimizing Large Language Model Inference with Cache Arbitration and Throttling
di: Zhou, Zhongchun, et al.
Pubblicazione: (2025)
di: Zhou, Zhongchun, et al.
Pubblicazione: (2025)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
di: Jo, Myeong Jun
Pubblicazione: (2026)
di: Jo, Myeong Jun
Pubblicazione: (2026)
Architecture-Aware LLM Inference Optimization on AMD Instinct GPUs: A Comprehensive Benchmark and Deployment Study
di: Georgiou, Athos
Pubblicazione: (2026)
di: Georgiou, Athos
Pubblicazione: (2026)
Exploring the Design Space for Message-Driven Systems for Dynamic Graph Processing using CCA
di: Chandio, Bibrak Qamar, et al.
Pubblicazione: (2024)
di: Chandio, Bibrak Qamar, et al.
Pubblicazione: (2024)
A Scalable NorthPole System with End-to-End Vertical Integration for Low-Latency and Energy-Efficient LLM Inference
di: DeBole, Michael V., et al.
Pubblicazione: (2025)
di: DeBole, Michael V., et al.
Pubblicazione: (2025)
Splitwise: Efficient generative LLM inference using phase splitting
di: Patel, Pratyush, et al.
Pubblicazione: (2023)
di: Patel, Pratyush, et al.
Pubblicazione: (2023)
MIMDRAM: An End-to-End Processing-Using-DRAM System for High-Throughput, Energy-Efficient and Programmer-Transparent Multiple-Instruction Multiple-Data Processing
di: Oliveira, Geraldo F., et al.
Pubblicazione: (2024)
di: Oliveira, Geraldo F., et al.
Pubblicazione: (2024)
Optimizing Communication for Latency Sensitive HPC Applications on up to 48 FPGAs Using ACCL
di: Meyer, Marius, et al.
Pubblicazione: (2024)
di: Meyer, Marius, et al.
Pubblicazione: (2024)
Efficient and Scalable Architecture for Multiple-chip Implementation of Simulated Bifurcation Machines
di: Kashimata, Tomoya, et al.
Pubblicazione: (2023)
di: Kashimata, Tomoya, et al.
Pubblicazione: (2023)
TeraNoC: A Multi-Channel 32-bit Fine-Grained, Hybrid Mesh-Crossbar NoC for Efficient Scale-up of 1000+ Core Shared-L1-Memory Clusters
di: Zhang, Yichao, et al.
Pubblicazione: (2025)
di: Zhang, Yichao, et al.
Pubblicazione: (2025)
Inside VOLT: Designing an Open-Source GPU Compiler
di: Jeong, Shinnung, et al.
Pubblicazione: (2025)
di: Jeong, Shinnung, et al.
Pubblicazione: (2025)
Compute Can't Handle the Truth: Why Communication Tax Prioritizes Memory and Interconnects in Modern AI Infrastructure
di: Jung, Myoungsoo
Pubblicazione: (2025)
di: Jung, Myoungsoo
Pubblicazione: (2025)
Exploring GPU-to-GPU Communication: Insights into Supercomputer Interconnects
di: De Sensi, Daniele, et al.
Pubblicazione: (2024)
di: De Sensi, Daniele, et al.
Pubblicazione: (2024)
CCSS: Hardware-Accelerated RTL Simulation with Fast Combinational Logic Computing and Sequential Logic Synchronization
di: Feng, Weigang, et al.
Pubblicazione: (2025)
di: Feng, Weigang, et al.
Pubblicazione: (2025)
LLAMP: Assessing Network Latency Tolerance of HPC Applications with Linear Programming
di: Shen, Siyuan, et al.
Pubblicazione: (2024)
di: Shen, Siyuan, et al.
Pubblicazione: (2024)
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
di: Ganjihal, Sanjeev Rao
Pubblicazione: (2026)
di: Ganjihal, Sanjeev Rao
Pubblicazione: (2026)
TAPA-CS: Enabling Scalable Accelerator Design on Distributed HBM-FPGAs
di: Prakriya, Neha, et al.
Pubblicazione: (2023)
di: Prakriya, Neha, et al.
Pubblicazione: (2023)
An Evaluation and Comparison of GPU Hardware and Solver Libraries for Accelerating the OPM Flow Reservoir Simulator
di: Qiu, Tong Dong, et al.
Pubblicazione: (2023)
di: Qiu, Tong Dong, et al.
Pubblicazione: (2023)
An End-to-End DNN Inference Framework for the SpiNNaker2 Neuromorphic MPSoC
di: Jobst, Matthias, et al.
Pubblicazione: (2025)
di: Jobst, Matthias, et al.
Pubblicazione: (2025)
A Reliable, Time-Predictable Heterogeneous SoC for AI-Enhanced Mixed-Criticality Edge Applications
di: Garofalo, Angelo, et al.
Pubblicazione: (2025)
di: Garofalo, Angelo, et al.
Pubblicazione: (2025)
Wattchmen: Watching the Wattchers -- High Fidelity, Flexible GPU Energy Modeling
di: Tran, Brandon, et al.
Pubblicazione: (2026)
di: Tran, Brandon, et al.
Pubblicazione: (2026)
Global Optimizations & Lightweight Dynamic Logic for Concurrency
di: Pati, Suchita, et al.
Pubblicazione: (2024)
di: Pati, Suchita, et al.
Pubblicazione: (2024)
GPU-Augmented OLAP Execution Engine: GPU Offloading
di: Chang, Ilsun
Pubblicazione: (2025)
di: Chang, Ilsun
Pubblicazione: (2025)
DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication
di: Pati, Suchita, et al.
Pubblicazione: (2025)
di: Pati, Suchita, et al.
Pubblicazione: (2025)
Documenti analoghi
-
TCDM Burst Access: Breaking the Bandwidth Barrier in Shared-L1 RVV Clusters Beyond 1000 FPUs
di: Shen, Diyou, et al.
Pubblicazione: (2025) -
Optimizing Foundation Model Inference on a Many-tiny-core Open-source RISC-V Platform
di: Potocnik, Viviane, et al.
Pubblicazione: (2024) -
MemPool Flavors: Between Versatility and Specialization in a RISC-V Manycore Cluster
di: Mazzola, Sergio, et al.
Pubblicazione: (2025) -
TeraPool: A Physical Design Aware, 1024 RISC-V Cores Shared-L1-Memory Scaled-up Cluster Design with High Bandwidth Main Memory Link
di: Zhang, Yichao, et al.
Pubblicazione: (2026) -
Lincoln AI Computing Survey (LAICS) and Trends
di: Reuther, Albert, et al.
Pubblicazione: (2025)