Fused Breadth-First Probabilistic Traversals on Distributed GPU Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Neff, Reece, Zarch, Mostafa Eghbali, Minutoli, Marco, Halappanavar, Mahantesh, Tumeo, Antonino, Kalyanaraman, Ananth, Becchi, Michela |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving the Efficiency of OpenCL Kernels through Pipes
by: Zarch, Mostafa Eghbali, et al.
Published: (2022)
by: Zarch, Mostafa Eghbali, et al.
Published: (2022)
GreediRIS: Scalable Influence Maximization using Distributed Streaming Maximum Cover
by: Barik, Reet, et al.
Published: (2024)
by: Barik, Reet, et al.
Published: (2024)
Picasso: Memory-Efficient Graph Coloring Using Palettes With Applications in Quantum Computing
by: Ferdous, S M, et al.
Published: (2024)
by: Ferdous, S M, et al.
Published: (2024)
Exploring Sparse Matrix Multiplication Kernels on the Cerebras CS-3
by: Shah, Milan, et al.
Published: (2026)
by: Shah, Milan, et al.
Published: (2026)
Anonymized Network Sensing using C++26 std::execution on GPUs
by: Mandulak, Michael, et al.
Published: (2025)
by: Mandulak, Michael, et al.
Published: (2025)
Performance-Driven Optimization of Parallel Breadth-First Search
by: Bhaskar, Marati, et al.
Published: (2025)
by: Bhaskar, Marati, et al.
Published: (2025)
DynLP: Parallel Dynamic Batch Update for Label Propagation in Semi-Supervised Learning
by: Shovan, S M, et al.
Published: (2026)
by: Shovan, S M, et al.
Published: (2026)
The Fused Kernel Library: A C++ API to Develop Highly-Efficient GPU Libraries
by: Amoros, Oscar, et al.
Published: (2025)
by: Amoros, Oscar, et al.
Published: (2025)
TurboFNO: High-Performance Fourier Neural Operator with Fused FFT-GEMM-iFFT on GPU
by: Wu, Shixun, et al.
Published: (2025)
by: Wu, Shixun, et al.
Published: (2025)
PULSE: Accelerating Distributed Pointer-Traversals on Disaggregated Memory (Extended Version)
by: Tang, Yupeng, et al.
Published: (2023)
by: Tang, Yupeng, et al.
Published: (2023)
Incidence Constraints in Hypergraph Partitioning on GPU
by: Ronzani, Marco, et al.
Published: (2026)
by: Ronzani, Marco, et al.
Published: (2026)
Towards Optimal Deterministic LOCAL Algorithms on Trees
by: Brandt, Sebastian, et al.
Published: (2025)
by: Brandt, Sebastian, et al.
Published: (2025)
Hypergraph Partitioning on GPU with Distinct Incident Hyperedges and Size Constraints
by: Ronzani, Marco, et al.
Published: (2026)
by: Ronzani, Marco, et al.
Published: (2026)
GPU-Accelerated Distributed QAOA on Large-scale HPC Ecosystems
by: Xu, Zhihao, et al.
Published: (2025)
by: Xu, Zhihao, et al.
Published: (2025)
SOLANET: Distributed Neighbor Graph Construction on GPU-Accelerated Systems
by: Iwabuchi, Keita, et al.
Published: (2026)
by: Iwabuchi, Keita, et al.
Published: (2026)
Proving Highly-Concurrent Traversals Correct
by: Feldman, Yotam M. Y., et al.
Published: (2020)
by: Feldman, Yotam M. Y., et al.
Published: (2020)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
by: Zhang, WenZheng, et al.
Published: (2024)
by: Zhang, WenZheng, et al.
Published: (2024)
Distributed OpenMP Offloading of OpenMC on Intel GPU MAX Accelerators
by: Fridman, Yehonatan, et al.
Published: (2024)
by: Fridman, Yehonatan, et al.
Published: (2024)
Parallel Collaborative ADMM Privacy Computing and Adaptive GPU Acceleration for Distributed Edge Networks
by: Xia, Mengchun, et al.
Published: (2026)
by: Xia, Mengchun, et al.
Published: (2026)
Characterizing Compute-Communication Overlap in GPU-Accelerated Distributed Deep Learning: Performance and Power Implications
by: Lee, Seonho, et al.
Published: (2025)
by: Lee, Seonho, et al.
Published: (2025)
Distributed Consensus Network: A Modularized Communication Framework and Reliability Probabilistic Analysis
by: Li, Yuetai, et al.
Published: (2025)
by: Li, Yuetai, et al.
Published: (2025)
PolyKAN: Efficient Fused GPU Operators for Polynomial Kolmogorov-Arnold Network Variants
by: Yu, Mingkun, et al.
Published: (2025)
by: Yu, Mingkun, et al.
Published: (2025)
An Online Probabilistic Distributed Tracing System
by: Toslali, M., et al.
Published: (2024)
by: Toslali, M., et al.
Published: (2024)
Federated Learning within Global Energy Budget over Heterogeneous Edge Accelerators
by: Banerjee, Roopkatha, et al.
Published: (2025)
by: Banerjee, Roopkatha, et al.
Published: (2025)
SIMPLE: Disaggregating Sampling from GPU Inference into a Decision Plane for Faster Distributed LLM Serving
by: Zhao, Bohan, et al.
Published: (2025)
by: Zhao, Bohan, et al.
Published: (2025)
EXaCTz: Guaranteed Extremum Graph and Contour Tree Preservation for Distributed- and GPU-Parallel Lossy Compression
by: Li, Yuxiao, et al.
Published: (2026)
by: Li, Yuxiao, et al.
Published: (2026)
SpecInF: Exploiting Idle GPU Resources in Distributed DL Training via Speculative Inference Filling
by: Lv, Cunchi, et al.
Published: (2025)
by: Lv, Cunchi, et al.
Published: (2025)
Improving GPU Multi-Tenancy Through Dynamic Multi-Instance GPU Reconfiguration
by: Wang, Tianyu, et al.
Published: (2024)
by: Wang, Tianyu, et al.
Published: (2024)
MixServe: An Automatic Distributed Serving System for MoE Models with Hybrid Parallelism Based on Fused Communication Algorithm
by: Zhou, Bowen, et al.
Published: (2026)
by: Zhou, Bowen, et al.
Published: (2026)
PRISM: Probabilistic Runtime Insights and Scalable Performance Modeling for Large-Scale Distributed Training
by: Golden, Alicia, et al.
Published: (2025)
by: Golden, Alicia, et al.
Published: (2025)
Traversal Learning: A Lossless And Efficient Distributed Learning Framework
by: Batbaatar, Erdenebileg, et al.
Published: (2025)
by: Batbaatar, Erdenebileg, et al.
Published: (2025)
GPU Under Pressure: Estimating Application's Stress via Telemetry and Performance Counters
by: Esposito, Giuseppe, et al.
Published: (2025)
by: Esposito, Giuseppe, et al.
Published: (2025)
Accelerating Biclique Counting on GPU
by: Qiu, Linshan, et al.
Published: (2024)
by: Qiu, Linshan, et al.
Published: (2024)
GPU Sharing with Triples Mode
by: Byun, Chansup, et al.
Published: (2024)
by: Byun, Chansup, et al.
Published: (2024)
ParvaGPU: Efficient Spatial GPU Sharing for Large-Scale DNN Inference in Cloud Environments
by: Lee, Munkyu, et al.
Published: (2024)
by: Lee, Munkyu, et al.
Published: (2024)
Accelerating Intra-Node GPU-to-GPU Communication Through Multi-Path Transfers with CUDA Graphs
by: Sojoodi, Amirhossein, et al.
Published: (2026)
by: Sojoodi, Amirhossein, et al.
Published: (2026)
Kubernetes in Action: Exploring the Performance of Kubernetes Distributions in the Cloud
by: Aqasizade, Hossein, et al.
Published: (2024)
by: Aqasizade, Hossein, et al.
Published: (2024)
DuaLip-GPU Technical Report
by: Dexter, Gregory, et al.
Published: (2026)
by: Dexter, Gregory, et al.
Published: (2026)
Predictable LLM Serving on GPU Clusters
by: Darzi, Erfan, et al.
Published: (2025)
by: Darzi, Erfan, et al.
Published: (2025)
GPU Accelerated Sparse Cholesky Factorization
by: Karsavuran, M. Ozan, et al.
Published: (2024)
by: Karsavuran, M. Ozan, et al.
Published: (2024)
Similar Items
-
Improving the Efficiency of OpenCL Kernels through Pipes
by: Zarch, Mostafa Eghbali, et al.
Published: (2022) -
GreediRIS: Scalable Influence Maximization using Distributed Streaming Maximum Cover
by: Barik, Reet, et al.
Published: (2024) -
Picasso: Memory-Efficient Graph Coloring Using Palettes With Applications in Quantum Computing
by: Ferdous, S M, et al.
Published: (2024) -
Exploring Sparse Matrix Multiplication Kernels on the Cerebras CS-3
by: Shah, Milan, et al.
Published: (2026) -
Anonymized Network Sensing using C++26 std::execution on GPUs
by: Mandulak, Michael, et al.
Published: (2025)