ARGO: An Auto-Tuning Runtime System for Scalable GNN Training on Multi-Core Processor
Fuente:
arXiv
Guardado en:
| Autores principales: | Lin, Yi-Chien, Chen, Yuyang, Gobriel, Sameh, Jain, Nilesh, Jha, Gopi Krishna, Prasanna, Viktor |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Unified CPU-GPU Protocol for GNN Training
por: Lin, Yi-Chien, et al.
Publicado: (2024)
por: Lin, Yi-Chien, et al.
Publicado: (2024)
ScalableHD: Scalable and High-Throughput Hyperdimensional Computing Inference on Multi-Core CPUs
por: Parikh, Dhruv, et al.
Publicado: (2025)
por: Parikh, Dhruv, et al.
Publicado: (2025)
GreenDyGNN: Runtime-Adaptive Energy-Efficient Communication for Distributed GNN Training
por: Niam, Arefin, et al.
Publicado: (2026)
por: Niam, Arefin, et al.
Publicado: (2026)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
por: Zhuang, Chen, et al.
Publicado: (2024)
por: Zhuang, Chen, et al.
Publicado: (2024)
RCOMPSs: A Scalable Runtime System for R Code Execution on Manycore Systems
por: Zhang, Xiran, et al.
Publicado: (2025)
por: Zhang, Xiran, et al.
Publicado: (2025)
PRISM: Probabilistic Runtime Insights and Scalable Performance Modeling for Large-Scale Distributed Training
por: Golden, Alicia, et al.
Publicado: (2025)
por: Golden, Alicia, et al.
Publicado: (2025)
Junctiond: Extending FaaS Runtimes with Kernel-Bypass
por: Saurez, Enrique, et al.
Publicado: (2024)
por: Saurez, Enrique, et al.
Publicado: (2024)
GraphLeap: Decoupling Graph Construction and Convolution for Vision GNN Acceleration on FPGA
por: Ramachandran, Anvitha, et al.
Publicado: (2026)
por: Ramachandran, Anvitha, et al.
Publicado: (2026)
Integrating and Characterizing HPC Task Runtime Systems for hybrid AI-HPC workloads
por: Merzky, Andre, et al.
Publicado: (2025)
por: Merzky, Andre, et al.
Publicado: (2025)
An AI-Native Runtime for Multi-Wearable Environments
por: Min, Chulhong, et al.
Publicado: (2024)
por: Min, Chulhong, et al.
Publicado: (2024)
Accelerating Sparse MTTKRP for Small Tensor Decomposition on GPU
por: Wijeratne, Sasindu, et al.
Publicado: (2025)
por: Wijeratne, Sasindu, et al.
Publicado: (2025)
AMPED: Accelerating MTTKRP for Billion-Scale Sparse Tensor Decomposition on Multiple GPUs
por: Wijeratne, Sasindu, et al.
Publicado: (2025)
por: Wijeratne, Sasindu, et al.
Publicado: (2025)
Accelerating Dynamic Image Graph Construction on FPGA for Vision GNNs
por: Ramachandran, Anvitha, et al.
Publicado: (2025)
por: Ramachandran, Anvitha, et al.
Publicado: (2025)
Scalable Runtime Architecture for Data-driven, Hybrid HPC and ML Workflow Applications
por: Merzky, Andre, et al.
Publicado: (2025)
por: Merzky, Andre, et al.
Publicado: (2025)
Ripple: Scalable Incremental GNN Inferencing on Large Streaming Graphs
por: Naman, Pranjal, et al.
Publicado: (2025)
por: Naman, Pranjal, et al.
Publicado: (2025)
Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training
por: Lu, Yishun, et al.
Publicado: (2026)
por: Lu, Yishun, et al.
Publicado: (2026)
HopGNN: Boosting Distributed GNN Training Efficiency via Feature-Centric Model Migration
por: Chen, Weijian, et al.
Publicado: (2024)
por: Chen, Weijian, et al.
Publicado: (2024)
Accelerating Mixed-Precision Out-of-Core Cholesky Factorization with Static Task Scheduling
por: Ren, Jie, et al.
Publicado: (2024)
por: Ren, Jie, et al.
Publicado: (2024)
Benchmarking the Performance of Large Language Models on the Cerebras Wafer Scale Engine
por: Zhang, Zuoning, et al.
Publicado: (2024)
por: Zhang, Zuoning, et al.
Publicado: (2024)
MQ-GNN: A Multi-Queue Pipelined Architecture for Scalable and Efficient GNN Training
por: Ullah, Irfan, et al.
Publicado: (2026)
por: Ullah, Irfan, et al.
Publicado: (2026)
Towards Affordable, Adaptive and Automatic GNN Training on CPU-GPU Heterogeneous Platforms
por: Qiao, Tong, et al.
Publicado: (2025)
por: Qiao, Tong, et al.
Publicado: (2025)
Understanding and Reducing Metadata-Driven Host Overheads in Sampling-Based GNN Training
por: Gong, Yidong, et al.
Publicado: (2026)
por: Gong, Yidong, et al.
Publicado: (2026)
Multi-level Memory-Centric Profiling on ARM Processors with ARM SPE
por: Miksits, Samuel, et al.
Publicado: (2024)
por: Miksits, Samuel, et al.
Publicado: (2024)
RapidGNN: Communication Efficient Large-Scale Distributed Training of Graph Neural Networks
por: Niam, Arefin, et al.
Publicado: (2025)
por: Niam, Arefin, et al.
Publicado: (2025)
CondenseGraph: Communication-Efficient Distributed GNN Training via On-the-Fly Graph Condensation
por: Zhang, Zizhao, et al.
Publicado: (2026)
por: Zhang, Zizhao, et al.
Publicado: (2026)
NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor Parallelism
por: Ai, Xin, et al.
Publicado: (2024)
por: Ai, Xin, et al.
Publicado: (2024)
LSM-GNN: Large-scale Storage-based Multi-GPU GNN Training by Optimizing Data Transfer Scheme
por: Park, Jeongmin Brian, et al.
Publicado: (2024)
por: Park, Jeongmin Brian, et al.
Publicado: (2024)
Accelerating Drug Discovery in AutoDock-GPU with Tensor Cores
por: Schieffer, Gabin, et al.
Publicado: (2024)
por: Schieffer, Gabin, et al.
Publicado: (2024)
AcOrch: Accelerating Sampling-based GNN Training under CPU-NPU Heterogeneous Environments
por: Chen, Kefu, et al.
Publicado: (2026)
por: Chen, Kefu, et al.
Publicado: (2026)
GriNNder: Breaking the Memory Capacity Wall in Full-Graph GNN Training with Storage Offloading
por: Song, Jaeyong, et al.
Publicado: (2026)
por: Song, Jaeyong, et al.
Publicado: (2026)
A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability
por: Liu, Ruitao, et al.
Publicado: (2026)
por: Liu, Ruitao, et al.
Publicado: (2026)
Auto-Tuning for OpenMP Dynamic Scheduling applied to Full Waveform Inversion
por: da Silva, Felipe H. S., et al.
Publicado: (2024)
por: da Silva, Felipe H. S., et al.
Publicado: (2024)
ACE-GNN: Adaptive GNN Co-Inference with System-Aware Scheduling in Dynamic Edge Environments
por: Zhou, Ao, et al.
Publicado: (2025)
por: Zhou, Ao, et al.
Publicado: (2025)
Co-designing a Programmable RISC-V Accelerator for MPC-based Energy and Thermal Management of Many-Core HPC Processors
por: Ottaviano, Alessandro, et al.
Publicado: (2025)
por: Ottaviano, Alessandro, et al.
Publicado: (2025)
HARP: A Taxonomy for Heterogeneous and Hierarchical Processors for Mixed-reuse Workloads
por: Garg, Raveesh, et al.
Publicado: (2025)
por: Garg, Raveesh, et al.
Publicado: (2025)
MuxTune: Efficient Multi-Task LLM Fine-Tuning in Multi-Tenant Datacenters via Spatial-Temporal Backbone Multiplexing
por: Xue, Chunyu, et al.
Publicado: (2026)
por: Xue, Chunyu, et al.
Publicado: (2026)
Cross-Layer Energy Analysis of Multimodal Training on Grace Hopper Superchips
por: Ahmed, Mahmoud, et al.
Publicado: (2026)
por: Ahmed, Mahmoud, et al.
Publicado: (2026)
Sparse MTTKRP Acceleration for Tensor Decomposition on GPU
por: Wijeratne, Sasindu, et al.
Publicado: (2024)
por: Wijeratne, Sasindu, et al.
Publicado: (2024)
Machine-Learning-Driven Runtime Optimization of BLAS Level 3 on Modern Multi-Core Systems
por: Xia, Yufan, et al.
Publicado: (2024)
por: Xia, Yufan, et al.
Publicado: (2024)
QEIL v2: Heterogeneous Computing for Edge Intelligence via Roofline-Derived Pareto-Optimal Energy Modeling and Multi-Objective Orchestration
por: Kumar, Satyam, et al.
Publicado: (2026)
por: Kumar, Satyam, et al.
Publicado: (2026)
Ejemplares similares
-
A Unified CPU-GPU Protocol for GNN Training
por: Lin, Yi-Chien, et al.
Publicado: (2024) -
ScalableHD: Scalable and High-Throughput Hyperdimensional Computing Inference on Multi-Core CPUs
por: Parikh, Dhruv, et al.
Publicado: (2025) -
GreenDyGNN: Runtime-Adaptive Energy-Efficient Communication for Distributed GNN Training
por: Niam, Arefin, et al.
Publicado: (2026) -
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
por: Zhuang, Chen, et al.
Publicado: (2024) -
RCOMPSs: A Scalable Runtime System for R Code Execution on Manycore Systems
por: Zhang, Xiran, et al.
Publicado: (2025)