GPU-centric Communication Schemes for HPC and ML Applications
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Namashivayam, Naveen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Application-Driven Exascale: The JUPITER Benchmark Suite
von: Herten, Andreas, et al.
Veröffentlicht: (2024)
von: Herten, Andreas, et al.
Veröffentlicht: (2024)
Aurora: Architecting Argonne's First Exascale Supercomputer for Accelerated Scientific Discovery
von: Allcock, William E., et al.
Veröffentlicht: (2025)
von: Allcock, William E., et al.
Veröffentlicht: (2025)
StreamFlow: cross-breeding cloud with HPC
von: Colonnelli, Iacopo, et al.
Veröffentlicht: (2020)
von: Colonnelli, Iacopo, et al.
Veröffentlicht: (2020)
Scalable Concurrent Queues for GPU
von: Shetty, Pratheek Prakash, et al.
Veröffentlicht: (2026)
von: Shetty, Pratheek Prakash, et al.
Veröffentlicht: (2026)
Automated Dynamic AI Inference Scaling on HPC-Infrastructure: Integrating Kubernetes, Slurm and vLLM
von: Trappen, Tim, et al.
Veröffentlicht: (2025)
von: Trappen, Tim, et al.
Veröffentlicht: (2025)
Exploring the Design Space for Message-Driven Systems for Dynamic Graph Processing using CCA
von: Chandio, Bibrak Qamar, et al.
Veröffentlicht: (2024)
von: Chandio, Bibrak Qamar, et al.
Veröffentlicht: (2024)
NM-SpMM: Accelerating Matrix Multiplication Using N:M Sparsity with GPGPU
von: Ma, Cong, et al.
Veröffentlicht: (2025)
von: Ma, Cong, et al.
Veröffentlicht: (2025)
PoCL-R: An Open Standard Based Offloading Layer for Heterogeneous Multi-Access Edge Computing with Server Side Scalability
von: Solanti, Jan, et al.
Veröffentlicht: (2023)
von: Solanti, Jan, et al.
Veröffentlicht: (2023)
Breaking (Global) Barriers in Parallel Stochastic Optimization with Wait-Avoiding Group Averaging
von: Li, Shigang, et al.
Veröffentlicht: (2020)
von: Li, Shigang, et al.
Veröffentlicht: (2020)
Scheduler-Driven Job Atomization
von: Konopa, Michal, et al.
Veröffentlicht: (2025)
von: Konopa, Michal, et al.
Veröffentlicht: (2025)
JASDA: Introducing Job-Aware Scheduling in Scheduler-Driven Job Atomization
von: Konopa, Michal, et al.
Veröffentlicht: (2025)
von: Konopa, Michal, et al.
Veröffentlicht: (2025)
Exploring GPU-to-GPU Communication: Insights into Supercomputer Interconnects
von: De Sensi, Daniele, et al.
Veröffentlicht: (2024)
von: De Sensi, Daniele, et al.
Veröffentlicht: (2024)
Bine Trees: Enhancing Collective Operations by Optimizing Communication Locality
von: De Sensi, Daniele, et al.
Veröffentlicht: (2025)
von: De Sensi, Daniele, et al.
Veröffentlicht: (2025)
Hybrid Quantum-HPC Middleware Systems for Adaptive Resource, Workload and Task Management
von: Mantha, Pradeep, et al.
Veröffentlicht: (2026)
von: Mantha, Pradeep, et al.
Veröffentlicht: (2026)
Speed, power and cost implications for GPU acceleration of Computational Fluid Dynamics on HPC systems
von: Cooper-Baldock, Zachary, et al.
Veröffentlicht: (2024)
von: Cooper-Baldock, Zachary, et al.
Veröffentlicht: (2024)
Stream parallel skeleton optimization
von: Aldinucci, Marco, et al.
Veröffentlicht: (2024)
von: Aldinucci, Marco, et al.
Veröffentlicht: (2024)
FLeNS: Federated Learning with Enhanced Nesterov-Newton Sketch
von: Gupta, Sunny, et al.
Veröffentlicht: (2024)
von: Gupta, Sunny, et al.
Veröffentlicht: (2024)
RACS-SADL: Robust and Understandable Randomized Consensus in the Cloud
von: Tennage, Pasindu, et al.
Veröffentlicht: (2024)
von: Tennage, Pasindu, et al.
Veröffentlicht: (2024)
Baxos: Backing off for Robust and Efficient Consensus
von: Tennage, Pasindu, et al.
Veröffentlicht: (2022)
von: Tennage, Pasindu, et al.
Veröffentlicht: (2022)
Toward a Universal GPU Instruction Set Architecture: A Cross-Vendor Analysis of Hardware-Invariant Computational Primitives in Parallel Processors
von: Abraham, Ojima, et al.
Veröffentlicht: (2026)
von: Abraham, Ojima, et al.
Veröffentlicht: (2026)
SpaDA: A Spatial Dataflow Architecture Programming Language
von: Gianinazzi, Lukas, et al.
Veröffentlicht: (2025)
von: Gianinazzi, Lukas, et al.
Veröffentlicht: (2025)
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation
von: Cheng, Long, et al.
Veröffentlicht: (2026)
von: Cheng, Long, et al.
Veröffentlicht: (2026)
Modular GPU Programming with Typed Perspectives
von: Bansal, Manya, et al.
Veröffentlicht: (2025)
von: Bansal, Manya, et al.
Veröffentlicht: (2025)
Construction of a Byzantine Linearizable SWMR Atomic Register from SWSR Atomic Registers
von: Kshemkalyani, Ajay D., et al.
Veröffentlicht: (2024)
von: Kshemkalyani, Ajay D., et al.
Veröffentlicht: (2024)
Using matrices in post-processing phase of CFD simulations
von: Argentini, Gianluca
Veröffentlicht: (2004)
von: Argentini, Gianluca
Veröffentlicht: (2004)
Rhizomes and Diffusions for Processing Highly Skewed Graphs on Fine-Grain Message-Driven Systems
von: Chandio, Bibrak Qamar, et al.
Veröffentlicht: (2024)
von: Chandio, Bibrak Qamar, et al.
Veröffentlicht: (2024)
Accelerating Matrix Multiplication: A Performance Comparison Between Multi-Core CPU and GPU
von: Ansari, Mufakir Qamar, et al.
Veröffentlicht: (2025)
von: Ansari, Mufakir Qamar, et al.
Veröffentlicht: (2025)
Relaxation for Efficient Asynchronous Queues
von: Baldwin, Samuel, et al.
Veröffentlicht: (2025)
von: Baldwin, Samuel, et al.
Veröffentlicht: (2025)
GPUnion: Autonomous GPU Sharing on Campus
von: Li, Yufang, et al.
Veröffentlicht: (2025)
von: Li, Yufang, et al.
Veröffentlicht: (2025)
PIM-STM: Software Transactional Memory for Processing-In-Memory Systems
von: Lopes, André, et al.
Veröffentlicht: (2024)
von: Lopes, André, et al.
Veröffentlicht: (2024)
Splitwise: Efficient generative LLM inference using phase splitting
von: Patel, Pratyush, et al.
Veröffentlicht: (2023)
von: Patel, Pratyush, et al.
Veröffentlicht: (2023)
Compute Can't Handle the Truth: Why Communication Tax Prioritizes Memory and Interconnects in Modern AI Infrastructure
von: Jung, Myoungsoo
Veröffentlicht: (2025)
von: Jung, Myoungsoo
Veröffentlicht: (2025)
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended
von: Kamath, Aditya K, et al.
Veröffentlicht: (2026)
von: Kamath, Aditya K, et al.
Veröffentlicht: (2026)
Fancy Some Chips for Your TeaStore? Modeling the Control of an Adaptable Discrete System
von: Gallone, Anna, et al.
Veröffentlicht: (2025)
von: Gallone, Anna, et al.
Veröffentlicht: (2025)
Feature-Aware Task-to-Core Allocation in Embedded Multi-core Platforms via Statistical Learning
von: Pivezhandi, Mohammad, et al.
Veröffentlicht: (2025)
von: Pivezhandi, Mohammad, et al.
Veröffentlicht: (2025)
Joint Training on AMD and NVIDIA GPUs
von: Hu, Jon, et al.
Veröffentlicht: (2026)
von: Hu, Jon, et al.
Veröffentlicht: (2026)
Laminar: A Probe-First Scheduling Paradigm with Deterministic Runtime Survival
von: Chu, Zhengyan
Veröffentlicht: (2026)
von: Chu, Zhengyan
Veröffentlicht: (2026)
Rank-Aware Resource Scheduling for Tightly-Coupled MPI Workloads on Kubernetes
von: Xie, Tianfang
Veröffentlicht: (2026)
von: Xie, Tianfang
Veröffentlicht: (2026)
MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices
von: Shakerdargah, Mohammadali, et al.
Veröffentlicht: (2024)
von: Shakerdargah, Mohammadali, et al.
Veröffentlicht: (2024)
LLM Agents for Interactive Workflow Provenance: Reference Architecture and Evaluation Methodology
von: Souza, Renan, et al.
Veröffentlicht: (2025)
von: Souza, Renan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Application-Driven Exascale: The JUPITER Benchmark Suite
von: Herten, Andreas, et al.
Veröffentlicht: (2024) -
Aurora: Architecting Argonne's First Exascale Supercomputer for Accelerated Scientific Discovery
von: Allcock, William E., et al.
Veröffentlicht: (2025) -
StreamFlow: cross-breeding cloud with HPC
von: Colonnelli, Iacopo, et al.
Veröffentlicht: (2020) -
Scalable Concurrent Queues for GPU
von: Shetty, Pratheek Prakash, et al.
Veröffentlicht: (2026) -
Automated Dynamic AI Inference Scaling on HPC-Infrastructure: Integrating Kubernetes, Slurm and vLLM
von: Trappen, Tim, et al.
Veröffentlicht: (2025)