Exploring the Design Space for Message-Driven Systems for Dynamic Graph Processing using CCA
Fuente:
arXiv
Salvato in:
| Autori principali: | Chandio, Bibrak Qamar, Brodowicz, Maciej, Sterling, Thomas |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Rhizomes and Diffusions for Processing Highly Skewed Graphs on Fine-Grain Message-Driven Systems
di: Chandio, Bibrak Qamar, et al.
Pubblicazione: (2024)
di: Chandio, Bibrak Qamar, et al.
Pubblicazione: (2024)
Structures and Techniques for Streaming Dynamic Graph Processing on Decentralized Message-Driven Systems
di: Chandio, Bibrak Qamar, et al.
Pubblicazione: (2024)
di: Chandio, Bibrak Qamar, et al.
Pubblicazione: (2024)
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation
di: Cheng, Long, et al.
Pubblicazione: (2026)
di: Cheng, Long, et al.
Pubblicazione: (2026)
Toward a Universal GPU Instruction Set Architecture: A Cross-Vendor Analysis of Hardware-Invariant Computational Primitives in Parallel Processors
di: Abraham, Ojima, et al.
Pubblicazione: (2026)
di: Abraham, Ojima, et al.
Pubblicazione: (2026)
pLUTo: Enabling Massively Parallel Computation in DRAM via Lookup Tables
di: Ferreira, João Dinis, et al.
Pubblicazione: (2021)
di: Ferreira, João Dinis, et al.
Pubblicazione: (2021)
A comprehensive evaluation of spatial co-execution on GPUs using MPS and MIG technologies
di: Villarrubia, Jorge, et al.
Pubblicazione: (2026)
di: Villarrubia, Jorge, et al.
Pubblicazione: (2026)
Exploring GPU-to-GPU Communication: Insights into Supercomputer Interconnects
di: De Sensi, Daniele, et al.
Pubblicazione: (2024)
di: De Sensi, Daniele, et al.
Pubblicazione: (2024)
LLaMCAT: Optimizing Large Language Model Inference with Cache Arbitration and Throttling
di: Zhou, Zhongchun, et al.
Pubblicazione: (2025)
di: Zhou, Zhongchun, et al.
Pubblicazione: (2025)
Application-Driven Exascale: The JUPITER Benchmark Suite
di: Herten, Andreas, et al.
Pubblicazione: (2024)
di: Herten, Andreas, et al.
Pubblicazione: (2024)
Global Optimizations & Lightweight Dynamic Logic for Concurrency
di: Pati, Suchita, et al.
Pubblicazione: (2024)
di: Pati, Suchita, et al.
Pubblicazione: (2024)
Lincoln AI Computing Survey (LAICS) and Trends
di: Reuther, Albert, et al.
Pubblicazione: (2025)
di: Reuther, Albert, et al.
Pubblicazione: (2025)
Aurora: Architecting Argonne's First Exascale Supercomputer for Accelerated Scientific Discovery
di: Allcock, William E., et al.
Pubblicazione: (2025)
di: Allcock, William E., et al.
Pubblicazione: (2025)
Wattchmen: Watching the Wattchers -- High Fidelity, Flexible GPU Energy Modeling
di: Tran, Brandon, et al.
Pubblicazione: (2026)
di: Tran, Brandon, et al.
Pubblicazione: (2026)
Athena: Synergizing Data Prefetching and Off-Chip Prediction via Online Reinforcement Learning
di: Bera, Rahul, et al.
Pubblicazione: (2026)
di: Bera, Rahul, et al.
Pubblicazione: (2026)
Machine Learning-Driven Intelligent Memory System Design: From On-Chip Caches to Storage
di: Bera, Rahul, et al.
Pubblicazione: (2026)
di: Bera, Rahul, et al.
Pubblicazione: (2026)
Mitigating the Memory Bottleneck with Machine Learning-Driven and Data-Aware Microarchitectural Techniques
di: Bera, Rahul
Pubblicazione: (2026)
di: Bera, Rahul
Pubblicazione: (2026)
Mestra: Exploring Migration on Virtualized CGRAs
di: Kyriazis, Agamemnon, et al.
Pubblicazione: (2026)
di: Kyriazis, Agamemnon, et al.
Pubblicazione: (2026)
T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives
di: Pati, Suchita, et al.
Pubblicazione: (2024)
di: Pati, Suchita, et al.
Pubblicazione: (2024)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
di: Jo, Myeong Jun
Pubblicazione: (2026)
di: Jo, Myeong Jun
Pubblicazione: (2026)
Efficient and Scalable Architecture for Multiple-chip Implementation of Simulated Bifurcation Machines
di: Kashimata, Tomoya, et al.
Pubblicazione: (2023)
di: Kashimata, Tomoya, et al.
Pubblicazione: (2023)
GPU-centric Communication Schemes for HPC and ML Applications
di: Namashivayam, Naveen
Pubblicazione: (2025)
di: Namashivayam, Naveen
Pubblicazione: (2025)
FPGA-Accelerated Lock Management and Transaction Processing: Architecture, Optimization, and Design Space Exploration
di: Zhu, Shien, et al.
Pubblicazione: (2026)
di: Zhu, Shien, et al.
Pubblicazione: (2026)
RACS-SADL: Robust and Understandable Randomized Consensus in the Cloud
di: Tennage, Pasindu, et al.
Pubblicazione: (2024)
di: Tennage, Pasindu, et al.
Pubblicazione: (2024)
Baxos: Backing off for Robust and Efficient Consensus
di: Tennage, Pasindu, et al.
Pubblicazione: (2022)
di: Tennage, Pasindu, et al.
Pubblicazione: (2022)
Compute Can't Handle the Truth: Why Communication Tax Prioritizes Memory and Interconnects in Modern AI Infrastructure
di: Jung, Myoungsoo
Pubblicazione: (2025)
di: Jung, Myoungsoo
Pubblicazione: (2025)
Inside VOLT: Designing an Open-Source GPU Compiler
di: Jeong, Shinnung, et al.
Pubblicazione: (2025)
di: Jeong, Shinnung, et al.
Pubblicazione: (2025)
Accelerating Precise End-to-End Simulation: Latency-Sensitive Many-core System Modeling
di: Li, Yinrong, et al.
Pubblicazione: (2026)
di: Li, Yinrong, et al.
Pubblicazione: (2026)
A WASM-Subset Stack Architecture for Low-cost FPGAs using Open-Source EDA Flows
di: Chakrabarti, Aradhya
Pubblicazione: (2025)
di: Chakrabarti, Aradhya
Pubblicazione: (2025)
Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures
di: Zhao, Chenggang, et al.
Pubblicazione: (2025)
di: Zhao, Chenggang, et al.
Pubblicazione: (2025)
Splitwise: Efficient generative LLM inference using phase splitting
di: Patel, Pratyush, et al.
Pubblicazione: (2023)
di: Patel, Pratyush, et al.
Pubblicazione: (2023)
A C++17 Thread Pool for High-Performance Scientific Computing
di: Shoshany, Barak
Pubblicazione: (2021)
di: Shoshany, Barak
Pubblicazione: (2021)
Design and Implementation of a RISC-V SoC with Custom DSP Accelerators for Edge Computing
di: Yadav, Priyanshu
Pubblicazione: (2025)
di: Yadav, Priyanshu
Pubblicazione: (2025)
Joint Training on AMD and NVIDIA GPUs
di: Hu, Jon, et al.
Pubblicazione: (2026)
di: Hu, Jon, et al.
Pubblicazione: (2026)
RAPID-Graph: Recursive All-Pairs Shortest Paths Using Processing-in-Memory for Dynamic Programming on Graphs
di: Chen, Yanru, et al.
Pubblicazione: (2025)
di: Chen, Yanru, et al.
Pubblicazione: (2025)
TriADA: Massively Parallel Trilinear Matrix-by-Tensor Multiply-Add Algorithm and Device Architecture for the Acceleration of 3D Discrete Transformations
di: Sedukhin, Stanislav, et al.
Pubblicazione: (2025)
di: Sedukhin, Stanislav, et al.
Pubblicazione: (2025)
Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs
di: Peng, Hongwu, et al.
Pubblicazione: (2023)
di: Peng, Hongwu, et al.
Pubblicazione: (2023)
Automated Dynamic AI Inference Scaling on HPC-Infrastructure: Integrating Kubernetes, Slurm and vLLM
di: Trappen, Tim, et al.
Pubblicazione: (2025)
di: Trappen, Tim, et al.
Pubblicazione: (2025)
SAKURAONE: Empowering Transparent and Open AI Platforms through Private-Sector HPC Investment in Japan
di: Konishi, Fumikazu
Pubblicazione: (2025)
di: Konishi, Fumikazu
Pubblicazione: (2025)
Optimizing Foundation Model Inference on a Many-tiny-core Open-source RISC-V Platform
di: Potocnik, Viviane, et al.
Pubblicazione: (2024)
di: Potocnik, Viviane, et al.
Pubblicazione: (2024)
FPGA-Accelerated RISC-V ISA Extensions for Efficient Neural Network Inference on Edge Devices
di: Parameshwara, Arya, et al.
Pubblicazione: (2025)
di: Parameshwara, Arya, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Rhizomes and Diffusions for Processing Highly Skewed Graphs on Fine-Grain Message-Driven Systems
di: Chandio, Bibrak Qamar, et al.
Pubblicazione: (2024) -
Structures and Techniques for Streaming Dynamic Graph Processing on Decentralized Message-Driven Systems
di: Chandio, Bibrak Qamar, et al.
Pubblicazione: (2024) -
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation
di: Cheng, Long, et al.
Pubblicazione: (2026) -
Toward a Universal GPU Instruction Set Architecture: A Cross-Vendor Analysis of Hardware-Invariant Computational Primitives in Parallel Processors
di: Abraham, Ojima, et al.
Pubblicazione: (2026) -
pLUTo: Enabling Massively Parallel Computation in DRAM via Lookup Tables
di: Ferreira, João Dinis, et al.
Pubblicazione: (2021)