Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhuang, Chen, Zhang, Lingqi, Wu, Du, Chen, Peng, Huang, Jiajun, Liu, Xin, Yokota, Rio, Dryden, Nikoli, Endo, Toshio, Matsuoka, Satoshi, Wahib, Mohamed |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SHIRO: Near-Optimal Communication Strategies for Distributed Sparse Matrix Multiplication
por: Zhuang, Chen, et al.
Publicado: (2025)
por: Zhuang, Chen, et al.
Publicado: (2025)
Can Tensor Cores Benefit Memory-Bound Kernels? (No!)
por: Zhang, Lingqi, et al.
Publicado: (2025)
por: Zhang, Lingqi, et al.
Publicado: (2025)
Extracting Practical, Actionable Energy Insights from Supercomputer Telemetry and Logs
por: Cornelius, Melanie, et al.
Publicado: (2025)
por: Cornelius, Melanie, et al.
Publicado: (2025)
Performance Evaluation of a Next-Generation SX-Aurora TSUBASA Vector Supercomputer
por: Takahashi, Keichi, et al.
Publicado: (2023)
por: Takahashi, Keichi, et al.
Publicado: (2023)
Lion Cub: Minimizing Communication Overhead in Distributed Lion
por: Ishikawa, Satoki, et al.
Publicado: (2024)
por: Ishikawa, Satoki, et al.
Publicado: (2024)
Evaluating HPC-Style CPU Performance and Cost in Virtualized Cloud Infrastructures
por: Tharwani, Jay, et al.
Publicado: (2025)
por: Tharwani, Jay, et al.
Publicado: (2025)
Dissecting the software-based measurement of CPU energy consumption: a comparative analysis
por: Raffin, Guillaume, et al.
Publicado: (2024)
por: Raffin, Guillaume, et al.
Publicado: (2024)
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
por: Wahlgren, Jacob, et al.
Publicado: (2025)
por: Wahlgren, Jacob, et al.
Publicado: (2025)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
por: Lin, Mao, et al.
Publicado: (2026)
por: Lin, Mao, et al.
Publicado: (2026)
Vectorization of Gradient Boosting of Decision Trees Prediction in the CatBoost Library for RISC-V Processors
por: Kozinov, Evgeny, et al.
Publicado: (2024)
por: Kozinov, Evgeny, et al.
Publicado: (2024)
MassiveGNN: Efficient Training via Prefetching for Massively Connected Distributed Graphs
por: Sarkar, Aishwarya, et al.
Publicado: (2024)
por: Sarkar, Aishwarya, et al.
Publicado: (2024)
xMem: A CPU-Based Approach for Accurate Estimation of GPU Memory in Deep Learning Training Workloads
por: Shi, Jiabo, et al.
Publicado: (2025)
por: Shi, Jiabo, et al.
Publicado: (2025)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
por: Karfakis, George, et al.
Publicado: (2025)
por: Karfakis, George, et al.
Publicado: (2025)
CPU-Limits kill Performance: Time to rethink Resource Control
por: Shetty, Chirag, et al.
Publicado: (2025)
por: Shetty, Chirag, et al.
Publicado: (2025)
Comparing CPU and GPU compute of PERMANOVA on MI300A
por: Sfiligoi, Igor
Publicado: (2025)
por: Sfiligoi, Igor
Publicado: (2025)
Performance Analysis of HPC applications on the Aurora Supercomputer: Exploring the Impact of HBM-Enabled Intel Xeon Max CPUs
por: Ibeid, Huda, et al.
Publicado: (2025)
por: Ibeid, Huda, et al.
Publicado: (2025)
Optimizing CPU Cache Utilization in Cloud VMs with Accurate Cache Abstraction
por: Tofigh, Mani, et al.
Publicado: (2025)
por: Tofigh, Mani, et al.
Publicado: (2025)
Accelerating Mobile Inference through Fine-Grained CPU-GPU Co-Execution
por: Li, Zhuojin, et al.
Publicado: (2025)
por: Li, Zhuojin, et al.
Publicado: (2025)
A Unified CPU-GPU Protocol for GNN Training
por: Lin, Yi-Chien, et al.
Publicado: (2024)
por: Lin, Yi-Chien, et al.
Publicado: (2024)
MQ-GNN: A Multi-Queue Pipelined Architecture for Scalable and Efficient GNN Training
por: Ullah, Irfan, et al.
Publicado: (2026)
por: Ullah, Irfan, et al.
Publicado: (2026)
Performance and scaling of the LFRic weather and climate model on different generations of HPE Cray EX supercomputers
por: Bull, J. Mark, et al.
Publicado: (2024)
por: Bull, J. Mark, et al.
Publicado: (2024)
A 4D Hybrid Algorithm to Scale Parallel Training to Thousands of GPUs
por: Singh, Siddharth, et al.
Publicado: (2023)
por: Singh, Siddharth, et al.
Publicado: (2023)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
por: Wang, Yuxin, et al.
Publicado: (2023)
por: Wang, Yuxin, et al.
Publicado: (2023)
Modeling the Impact of Fiber Latency on Compute-Communication Overlap in Geo-Distributed Multi-Datacenter AI Training
por: Papavasileiou, Ioannis, et al.
Publicado: (2026)
por: Papavasileiou, Ioannis, et al.
Publicado: (2026)
Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
por: Titopoulos, Vasileios, et al.
Publicado: (2025)
por: Titopoulos, Vasileios, et al.
Publicado: (2025)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
por: Zhang, Li, et al.
Publicado: (2025)
por: Zhang, Li, et al.
Publicado: (2025)
Performance Optimization in Stream Processing Systems: Experiment-Driven Configuration Tuning for Kafka Streams
por: Chen, David, et al.
Publicado: (2026)
por: Chen, David, et al.
Publicado: (2026)
AcOrch: Accelerating Sampling-based GNN Training under CPU-NPU Heterogeneous Environments
por: Chen, Kefu, et al.
Publicado: (2026)
por: Chen, Kefu, et al.
Publicado: (2026)
Preliminary report: Initial evaluation of StdPar implementations on AMD GPUs for HPC
por: Lin, Wei-Chen, et al.
Publicado: (2024)
por: Lin, Wei-Chen, et al.
Publicado: (2024)
InkStream: Real-time GNN Inference on Streaming Graphs via Incremental Update
por: Wu, Dan, et al.
Publicado: (2023)
por: Wu, Dan, et al.
Publicado: (2023)
Recorder: Comprehensive Parallel I/O Tracing and Analysis
por: Wang, Chen, et al.
Publicado: (2025)
por: Wang, Chen, et al.
Publicado: (2025)
Understanding Large-Scale Plasma Simulation Challenges for Fusion Energy on Supercomputers
por: Williams, Jeremy J., et al.
Publicado: (2024)
por: Williams, Jeremy J., et al.
Publicado: (2024)
Parallel I/O Characterization and Optimization on Large-Scale HPC Systems: A 360-Degree Survey
por: Ather, Hammad, et al.
Publicado: (2024)
por: Ather, Hammad, et al.
Publicado: (2024)
Orthrus: Accelerating Multi-BFT Consensus through Concurrent Partial Ordering of Transactions (Extended Version)
por: Lyu, Hanzheng, et al.
Publicado: (2024)
por: Lyu, Hanzheng, et al.
Publicado: (2024)
Towards Affordable, Adaptive and Automatic GNN Training on CPU-GPU Heterogeneous Platforms
por: Qiao, Tong, et al.
Publicado: (2025)
por: Qiao, Tong, et al.
Publicado: (2025)
ARGO: An Auto-Tuning Runtime System for Scalable GNN Training on Multi-Core Processor
por: Lin, Yi-Chien, et al.
Publicado: (2024)
por: Lin, Yi-Chien, et al.
Publicado: (2024)
FalconFS: Distributed File System for Large-Scale Deep Learning Pipeline
por: Xu, Jingwei, et al.
Publicado: (2025)
por: Xu, Jingwei, et al.
Publicado: (2025)
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
por: Wang, Yuxin, et al.
Publicado: (2024)
por: Wang, Yuxin, et al.
Publicado: (2024)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
por: Vellaisamy, Prabhu, et al.
Publicado: (2025)
por: Vellaisamy, Prabhu, et al.
Publicado: (2025)
Massimult: A Novel Parallel CPU Architecture Based on Combinator Reduction
por: Nicklisch-Franken, Jurgen, et al.
Publicado: (2024)
por: Nicklisch-Franken, Jurgen, et al.
Publicado: (2024)
Ejemplares similares
-
SHIRO: Near-Optimal Communication Strategies for Distributed Sparse Matrix Multiplication
por: Zhuang, Chen, et al.
Publicado: (2025) -
Can Tensor Cores Benefit Memory-Bound Kernels? (No!)
por: Zhang, Lingqi, et al.
Publicado: (2025) -
Extracting Practical, Actionable Energy Insights from Supercomputer Telemetry and Logs
por: Cornelius, Melanie, et al.
Publicado: (2025) -
Performance Evaluation of a Next-Generation SX-Aurora TSUBASA Vector Supercomputer
por: Takahashi, Keichi, et al.
Publicado: (2023) -
Lion Cub: Minimizing Communication Overhead in Distributed Lion
por: Ishikawa, Satoki, et al.
Publicado: (2024)