Communication-free Sampling and 4D Hybrid Parallelism for Scalable Mini-batch GNN Training
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Cunyang, Singh, Siddharth, Sarkar, Aishwarya, Nichols, Daniel, Patel, Tisha, Ranjan, Aditya K., Ghosh, Sayan, Jannesari, Ali, Tallent, Nathan R., Bhatele, Abhinav |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Plexus: Taming Billion-edge Graphs with 3D Parallel Full-graph GNN Training
by: Ranjan, Aditya K., et al.
Published: (2025)
by: Ranjan, Aditya K., et al.
Published: (2025)
MassiveGNN: Efficient Training via Prefetching for Massively Connected Distributed Graphs
by: Sarkar, Aishwarya, et al.
Published: (2024)
by: Sarkar, Aishwarya, et al.
Published: (2024)
NOMAD: Generating Embeddings for Massive Distributed Graphs
by: Sarkar, Aishwarya, et al.
Published: (2026)
by: Sarkar, Aishwarya, et al.
Published: (2026)
Rudder: Steering Prefetching in Distributed GNN Training using LLM Agents
by: Sarkar, Aishwarya, et al.
Published: (2026)
by: Sarkar, Aishwarya, et al.
Published: (2026)
The Big Send-off: Scalable and Performant Collectives for Deep Learning
by: Singh, Siddharth, et al.
Published: (2025)
by: Singh, Siddharth, et al.
Published: (2025)
A 4D Hybrid Algorithm to Scale Parallel Training to Thousands of GPUs
by: Singh, Siddharth, et al.
Published: (2023)
by: Singh, Siddharth, et al.
Published: (2023)
HPC-Coder-V2: Studying Code LLMs Across Low-Resource Parallel Languages
by: Chaturvedi, Aman, et al.
Published: (2024)
by: Chaturvedi, Aman, et al.
Published: (2024)
Automated Programmatic Performance Analysis of Parallel Programs
by: Cankur, Onur, et al.
Published: (2024)
by: Cankur, Onur, et al.
Published: (2024)
HPC-Coder: Modeling Parallel Programs using Large Language Models
by: Nichols, Daniel, et al.
Published: (2023)
by: Nichols, Daniel, et al.
Published: (2023)
Can Large Language Models Write Parallel Code?
by: Nichols, Daniel, et al.
Published: (2024)
by: Nichols, Daniel, et al.
Published: (2024)
Pipit: Scripting the analysis of parallel execution traces
by: Bhatele, Abhinav, et al.
Published: (2023)
by: Bhatele, Abhinav, et al.
Published: (2023)
ParEval-Repo: A Benchmark Suite for Evaluating LLMs with Repository-level HPC Translation Tasks
by: Davis, Joshua H., et al.
Published: (2025)
by: Davis, Joshua H., et al.
Published: (2025)
Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers
by: Singh, Siddharth, et al.
Published: (2025)
by: Singh, Siddharth, et al.
Published: (2025)
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
by: Nichols, Daniel, et al.
Published: (2025)
by: Nichols, Daniel, et al.
Published: (2025)
Analytics of Longitudinal System Monitoring Data for Performance Prediction
by: Costello, Ian J., et al.
Published: (2020)
by: Costello, Ian J., et al.
Published: (2020)
Optimizing Agentic Language Model Inference via Speculative Tool Calls
by: Nichols, Daniel, et al.
Published: (2025)
by: Nichols, Daniel, et al.
Published: (2025)
CARAT: Client-Side Adaptive RPC and Cache Co-Tuning for Parallel File Systems
by: Rashid, Md Hasanur, et al.
Published: (2026)
by: Rashid, Md Hasanur, et al.
Published: (2026)
Performance-Aligned LLMs for Generating Fast Code
by: Nichols, Daniel, et al.
Published: (2024)
by: Nichols, Daniel, et al.
Published: (2024)
Characterizing Production GPU Workloads using System-wide Telemetry Data
by: Cankur, Onur, et al.
Published: (2025)
by: Cankur, Onur, et al.
Published: (2025)
HydroGAT: Distributed Heterogeneous Graph Attention Transformer for Spatiotemporal Flood Prediction
by: Sarkar, Aishwarya, et al.
Published: (2025)
by: Sarkar, Aishwarya, et al.
Published: (2025)
DistShap: Scalable GNN Explanations with Distributed Shapley Values
by: Akkas, Selahattin, et al.
Published: (2025)
by: Akkas, Selahattin, et al.
Published: (2025)
KEET: Explaining Performance of GPU Kernels Using LLM Agents
by: Davis, Joshua H., et al.
Published: (2026)
by: Davis, Joshua H., et al.
Published: (2026)
ML-based Modeling to Predict I/O Performance on Different Storage Sub-systems
by: Xu, Yiheng, et al.
Published: (2023)
by: Xu, Yiheng, et al.
Published: (2023)
On The Reproducibility Limitations of RAG Systems
by: Wang, Baiqiang, et al.
Published: (2025)
by: Wang, Baiqiang, et al.
Published: (2025)
BigSUMO: A Scalable Framework for Big Data Traffic Analytics and Parallel Simulation
by: Sengupta, Rahul, et al.
Published: (2026)
by: Sengupta, Rahul, et al.
Published: (2026)
Understanding and Improving Communication Performance in Multi-node LLM Inference
by: Singhania, Prajwal, et al.
Published: (2025)
by: Singhania, Prajwal, et al.
Published: (2025)
PowerTrip: Exploiting Federated Heterogeneous Datacenter Power for Distributed ML Training
by: Mehboob, Talha, et al.
Published: (2025)
by: Mehboob, Talha, et al.
Published: (2025)
MPI Errors Detection using GNN Embedding and Vector Embedding over LLVM IR
by: Karchi, Jad El, et al.
Published: (2024)
by: Karchi, Jad El, et al.
Published: (2024)
ElasWave: An Elastic-Native System for Scalable Hybrid-Parallel Training
by: Kang, Xueze, et al.
Published: (2025)
by: Kang, Xueze, et al.
Published: (2025)
Ripple: Scalable Incremental GNN Inferencing on Large Streaming Graphs
by: Naman, Pranjal, et al.
Published: (2025)
by: Naman, Pranjal, et al.
Published: (2025)
LLMTailor: A Layer-wise Tailoring Tool for Efficient Checkpointing of Large Language Models
by: Sun, Minqiu, et al.
Published: (2026)
by: Sun, Minqiu, et al.
Published: (2026)
NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor Parallelism
by: Ai, Xin, et al.
Published: (2024)
by: Ai, Xin, et al.
Published: (2024)
Overcoming Memory Constraints in Quantum Circuit Simulation with a High-Fidelity Compression Framework
by: Zhang, Boyuan, et al.
Published: (2024)
by: Zhang, Boyuan, et al.
Published: (2024)
Dynamic Detection of Inefficient Data Mapping Patterns in Heterogeneous OpenMP Applications
by: Marzen, Luke, et al.
Published: (2026)
by: Marzen, Luke, et al.
Published: (2026)
Static Generation of Efficient OpenMP Offload Data Mappings
by: Marzen, Luke, et al.
Published: (2024)
by: Marzen, Luke, et al.
Published: (2024)
ARGO: An Auto-Tuning Runtime System for Scalable GNN Training on Multi-Core Processor
by: Lin, Yi-Chien, et al.
Published: (2024)
by: Lin, Yi-Chien, et al.
Published: (2024)
QoSFlow: Ensuring Service Quality of Distributed Workflows Using Interpretable Sensitivity Models
by: Rashid, Md Hasanur, et al.
Published: (2026)
by: Rashid, Md Hasanur, et al.
Published: (2026)
MIREncoder: Multi-modal IR-based Pretrained Embeddings for Performance Optimizations
by: Dutta, Akash, et al.
Published: (2024)
by: Dutta, Akash, et al.
Published: (2024)
An Analytical Overview Of Virtual Machine Load Balancing Scheduling Algorithms with their Comparative Case Study
by: Vaidya, Priyank, et al.
Published: (2025)
by: Vaidya, Priyank, et al.
Published: (2025)
Workflow Mini-Apps: Portable, Scalable, Tunable & Faithful Representations of Scientific Workflows
by: Kilic, Ozgur Ozan, et al.
Published: (2024)
by: Kilic, Ozgur Ozan, et al.
Published: (2024)
Similar Items
-
Plexus: Taming Billion-edge Graphs with 3D Parallel Full-graph GNN Training
by: Ranjan, Aditya K., et al.
Published: (2025) -
MassiveGNN: Efficient Training via Prefetching for Massively Connected Distributed Graphs
by: Sarkar, Aishwarya, et al.
Published: (2024) -
NOMAD: Generating Embeddings for Massive Distributed Graphs
by: Sarkar, Aishwarya, et al.
Published: (2026) -
Rudder: Steering Prefetching in Distributed GNN Training using LLM Agents
by: Sarkar, Aishwarya, et al.
Published: (2026) -
The Big Send-off: Scalable and Performant Collectives for Deep Learning
by: Singh, Siddharth, et al.
Published: (2025)