How Fast Can Graph Computations Go on Fine-grained Parallel Architectures
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Yuqing, Colley, Charles, Wheatman, Brian, Su, Jiya, Gleich, David F., Chien, Andrew A. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Modernizing Amdahl's Law: How AI Scaling Laws Shape Computer Architecture
di: Lu, Chien-Ping
Pubblicazione: (2026)
di: Lu, Chien-Ping
Pubblicazione: (2026)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
di: Zhang, Chen, et al.
Pubblicazione: (2026)
di: Zhang, Chen, et al.
Pubblicazione: (2026)
Deep Learning and Machine Learning with GPGPU and CUDA: Unlocking the Power of Parallel Computing
di: Li, Ming, et al.
Pubblicazione: (2024)
di: Li, Ming, et al.
Pubblicazione: (2024)
CCSS: Hardware-Accelerated RTL Simulation with Fast Combinational Logic Computing and Sequential Logic Synchronization
di: Feng, Weigang, et al.
Pubblicazione: (2025)
di: Feng, Weigang, et al.
Pubblicazione: (2025)
Parendi: Thousand-Way Parallel RTL Simulation
di: Emami, Mahyar, et al.
Pubblicazione: (2024)
di: Emami, Mahyar, et al.
Pubblicazione: (2024)
Datapath Combinational Equivalence Checking With Hybrid Sweeping Engines and Parallelization
di: Chen, Zhihan, et al.
Pubblicazione: (2024)
di: Chen, Zhihan, et al.
Pubblicazione: (2024)
General-Purpose Multicore Architectures
di: Ghose, Saugata
Pubblicazione: (2024)
di: Ghose, Saugata
Pubblicazione: (2024)
Dynamic Simultaneous Multithreaded Architecture
di: Ortiz-Arroyo, Daniel, et al.
Pubblicazione: (2024)
di: Ortiz-Arroyo, Daniel, et al.
Pubblicazione: (2024)
PIUMA: Programmable Integrated Unified Memory Architecture
di: Aananthakrishnan, Sriram, et al.
Pubblicazione: (2020)
di: Aananthakrishnan, Sriram, et al.
Pubblicazione: (2020)
Memory-Centric Computing: Solving Computing's Memory Problem
di: Mutlu, Onur, et al.
Pubblicazione: (2025)
di: Mutlu, Onur, et al.
Pubblicazione: (2025)
SpArch: Efficient Architecture for Sparse Matrix Multiplication
di: Zhang, Zhekai, et al.
Pubblicazione: (2020)
di: Zhang, Zhekai, et al.
Pubblicazione: (2020)
Efficient Architecture for RISC-V Vector Memory Access
di: Guan, Hongyi, et al.
Pubblicazione: (2025)
di: Guan, Hongyi, et al.
Pubblicazione: (2025)
Taming Offload Overheads in a Massively Parallel Open-Source RISC-V MPSoC: Analysis and Optimization
di: Colagrande, Luca, et al.
Pubblicazione: (2025)
di: Colagrande, Luca, et al.
Pubblicazione: (2025)
Navigating the Landscape of Distributed File Systems: Architectures, Implementations, and Considerations
di: Pan, Xueting, et al.
Pubblicazione: (2024)
di: Pan, Xueting, et al.
Pubblicazione: (2024)
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
di: li, Fei, et al.
Pubblicazione: (2026)
di: li, Fei, et al.
Pubblicazione: (2026)
DCRA: A Distributed Chiplet-based Reconfigurable Architecture for Irregular Applications
di: Orenes-Vera, Marcelo, et al.
Pubblicazione: (2023)
di: Orenes-Vera, Marcelo, et al.
Pubblicazione: (2023)
iHAC: A Hybrid Cluster Architecture for Enhanced Performance and Resilience
di: Muntaka, Siddique Abubakr, et al.
Pubblicazione: (2026)
di: Muntaka, Siddique Abubakr, et al.
Pubblicazione: (2026)
A Heterogeneous Chiplet Architecture for Accelerating End-to-End Transformer Models
di: Sharma, Harsh, et al.
Pubblicazione: (2023)
di: Sharma, Harsh, et al.
Pubblicazione: (2023)
Microbenchmark-Driven Analytical Performance Modeling Across Modern GPU Architectures
di: Jarmusch, Aaron, et al.
Pubblicazione: (2026)
di: Jarmusch, Aaron, et al.
Pubblicazione: (2026)
New Tools, Programming Models, and System Support for Processing-in-Memory Architectures
di: Oliveira, Geraldo F.
Pubblicazione: (2025)
di: Oliveira, Geraldo F.
Pubblicazione: (2025)
FinGraV: Methodology for Fine-Grain GPU Power Visibility and Insights
di: Singhania, Varsha, et al.
Pubblicazione: (2024)
di: Singhania, Varsha, et al.
Pubblicazione: (2024)
Fine-Grained Power and Energy Attribution on AMD GPU/APU-Based Exascale Nodes
di: McDaniel, Adam, et al.
Pubblicazione: (2026)
di: McDaniel, Adam, et al.
Pubblicazione: (2026)
Exploring the Efficiency of 3D-Stacked AI Chip Architecture for LLM Inference with Voxel
di: Liu, Yiqi, et al.
Pubblicazione: (2026)
di: Liu, Yiqi, et al.
Pubblicazione: (2026)
RAPID-Graph: Recursive All-Pairs Shortest Paths Using Processing-in-Memory for Dynamic Programming on Graphs
di: Chen, Yanru, et al.
Pubblicazione: (2025)
di: Chen, Yanru, et al.
Pubblicazione: (2025)
PID-Comm: A Fast and Flexible Collective Communication Framework for Commodity Processing-in-DIMM Devices
di: Noh, Si Ung, et al.
Pubblicazione: (2024)
di: Noh, Si Ung, et al.
Pubblicazione: (2024)
HgPCN: A Heterogeneous Architecture for E2E Embedded Point Cloud Inference
di: Gao, Yiming, et al.
Pubblicazione: (2025)
di: Gao, Yiming, et al.
Pubblicazione: (2025)
A Survey of Real-time Scheduling on Accelerator-based Heterogeneous Architecture for Time Critical Applications
di: Zou, An, et al.
Pubblicazione: (2025)
di: Zou, An, et al.
Pubblicazione: (2025)
Revisiting Computational Storage for Data Integrity and Security
di: Shi, Chao, et al.
Pubblicazione: (2025)
di: Shi, Chao, et al.
Pubblicazione: (2025)
GigaAPI for GPU Parallelization
di: Suvarna, M., et al.
Pubblicazione: (2025)
di: Suvarna, M., et al.
Pubblicazione: (2025)
Optimizing Task Scheduling in Fog Computing with Deadline Awareness
di: Sirjani, Mohammad Sadegh, et al.
Pubblicazione: (2025)
di: Sirjani, Mohammad Sadegh, et al.
Pubblicazione: (2025)
Memory-Centric Computing: Recent Advances in Processing-in-DRAM
di: Mutlu, Onur, et al.
Pubblicazione: (2024)
di: Mutlu, Onur, et al.
Pubblicazione: (2024)
DeFiNES: Enabling Fast Exploration of the Depth-first Scheduling Space for DNN Accelerators through Analytical Modeling
di: Mei, Linyan, et al.
Pubblicazione: (2022)
di: Mei, Linyan, et al.
Pubblicazione: (2022)
Optimizing Distributed ML Communication with Fused Computation-Collective Operations
di: Punniyamurthy, Kishore, et al.
Pubblicazione: (2023)
di: Punniyamurthy, Kishore, et al.
Pubblicazione: (2023)
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
di: Agrawal, Anirudha, et al.
Pubblicazione: (2024)
di: Agrawal, Anirudha, et al.
Pubblicazione: (2024)
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
di: Zhang, Qijun, et al.
Pubblicazione: (2026)
di: Zhang, Qijun, et al.
Pubblicazione: (2026)
XDMA: A Distributed, Extensible DMA Architecture for Layout-Flexible Data Movements in Heterogeneous Multi-Accelerator SoCs
di: Kong, Fanchen, et al.
Pubblicazione: (2025)
di: Kong, Fanchen, et al.
Pubblicazione: (2025)
Parallelizing a modern GPU simulator
di: Huerta, Rodrigo, et al.
Pubblicazione: (2025)
di: Huerta, Rodrigo, et al.
Pubblicazione: (2025)
Knowledge-Guided Attention-Inspired Learning for Task Offloading in Vehicle Edge Computing
di: Ma, Ke, et al.
Pubblicazione: (2025)
di: Ma, Ke, et al.
Pubblicazione: (2025)
BlockAMC: Scalable In-Memory Analog Matrix Computing for Solving Linear Systems
di: Pan, Lunshuai, et al.
Pubblicazione: (2024)
di: Pan, Lunshuai, et al.
Pubblicazione: (2024)
UniFormer: Unified and Efficient Transformer for Reasoning Across General and Custom Computing
di: Ran, Zhuoheng, et al.
Pubblicazione: (2025)
di: Ran, Zhuoheng, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Modernizing Amdahl's Law: How AI Scaling Laws Shape Computer Architecture
di: Lu, Chien-Ping
Pubblicazione: (2026) -
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
di: Zhang, Chen, et al.
Pubblicazione: (2026) -
Deep Learning and Machine Learning with GPGPU and CUDA: Unlocking the Power of Parallel Computing
di: Li, Ming, et al.
Pubblicazione: (2024) -
CCSS: Hardware-Accelerated RTL Simulation with Fast Combinational Logic Computing and Sequential Logic Synchronization
di: Feng, Weigang, et al.
Pubblicazione: (2025) -
Parendi: Thousand-Way Parallel RTL Simulation
di: Emami, Mahyar, et al.
Pubblicazione: (2024)