MassiveGNN: Efficient Training via Prefetching for Massively Connected Distributed Graphs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sarkar, Aishwarya, Ghosh, Sayan, Tallent, Nathan R., Jannesari, Ali |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Rudder: Steering Prefetching in Distributed GNN Training using LLM Agents
von: Sarkar, Aishwarya, et al.
Veröffentlicht: (2026)
von: Sarkar, Aishwarya, et al.
Veröffentlicht: (2026)
NOMAD: Generating Embeddings for Massive Distributed Graphs
von: Sarkar, Aishwarya, et al.
Veröffentlicht: (2026)
von: Sarkar, Aishwarya, et al.
Veröffentlicht: (2026)
Communication-free Sampling and 4D Hybrid Parallelism for Scalable Mini-batch GNN Training
von: Wei, Cunyang, et al.
Veröffentlicht: (2026)
von: Wei, Cunyang, et al.
Veröffentlicht: (2026)
MIREncoder: Multi-modal IR-based Pretrained Embeddings for Performance Optimizations
von: Dutta, Akash, et al.
Veröffentlicht: (2024)
von: Dutta, Akash, et al.
Veröffentlicht: (2024)
InkStream: Real-time GNN Inference on Streaming Graphs via Incremental Update
von: Wu, Dan, et al.
Veröffentlicht: (2023)
von: Wu, Dan, et al.
Veröffentlicht: (2023)
MQ-GNN: A Multi-Queue Pipelined Architecture for Scalable and Efficient GNN Training
von: Ullah, Irfan, et al.
Veröffentlicht: (2026)
von: Ullah, Irfan, et al.
Veröffentlicht: (2026)
Distributed Matrix-Based Sampling for Graph Neural Network Training
von: Tripathy, Alok, et al.
Veröffentlicht: (2023)
von: Tripathy, Alok, et al.
Veröffentlicht: (2023)
QoSFlow: Ensuring Service Quality of Distributed Workflows Using Interpretable Sensitivity Models
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
HydroGAT: Distributed Heterogeneous Graph Attention Transformer for Spatiotemporal Flood Prediction
von: Sarkar, Aishwarya, et al.
Veröffentlicht: (2025)
von: Sarkar, Aishwarya, et al.
Veröffentlicht: (2025)
CARAT: Client-Side Adaptive RPC and Cache Co-Tuning for Parallel File Systems
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
ProTrain: Efficient LLM Training via Memory-Aware Techniques
von: Yang, Hanmei, et al.
Veröffentlicht: (2024)
von: Yang, Hanmei, et al.
Veröffentlicht: (2024)
KVDirect: Distributed Disaggregated LLM Inference
von: Chen, Shiyang, et al.
Veröffentlicht: (2024)
von: Chen, Shiyang, et al.
Veröffentlicht: (2024)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
Towards Universal Performance Modeling for Machine Learning Training on Multi-GPU Platforms
von: Lin, Zhongyi, et al.
Veröffentlicht: (2024)
von: Lin, Zhongyi, et al.
Veröffentlicht: (2024)
AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence Parallelism
von: Gupta, Ahan, et al.
Veröffentlicht: (2026)
von: Gupta, Ahan, et al.
Veröffentlicht: (2026)
When Less is More: Achieving Faster Convergence in Distributed Edge Machine Learning
von: Basani, Advik Raj, et al.
Veröffentlicht: (2024)
von: Basani, Advik Raj, et al.
Veröffentlicht: (2024)
Training Time Prediction for Mixed Precision-based Distributed Training
von: Kang, Minchul, et al.
Veröffentlicht: (2026)
von: Kang, Minchul, et al.
Veröffentlicht: (2026)
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models
von: Ding, Shiwei, et al.
Veröffentlicht: (2025)
von: Ding, Shiwei, et al.
Veröffentlicht: (2025)
xMem: A CPU-Based Approach for Accurate Estimation of GPU Memory in Deep Learning Training Workloads
von: Shi, Jiabo, et al.
Veröffentlicht: (2025)
von: Shi, Jiabo, et al.
Veröffentlicht: (2025)
ReLATE: Learning Efficient Sparse Encoding for High-Performance Tensor Decomposition
von: Helal, Ahmed E., et al.
Veröffentlicht: (2025)
von: Helal, Ahmed E., et al.
Veröffentlicht: (2025)
AutoChunk: Automated Activation Chunk for Memory-Efficient Long Sequence Inference
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
iSpLib: A Library for Accelerating Graph Neural Networks using Auto-tuned Sparse Operations
von: Anik, Md Saidul Hoque, et al.
Veröffentlicht: (2024)
von: Anik, Md Saidul Hoque, et al.
Veröffentlicht: (2024)
Cost-Efficient LLM Training with Lifetime-Aware Tensor Offloading via GPUDirect Storage
von: Yuan, Ziqi, et al.
Veröffentlicht: (2025)
von: Yuan, Ziqi, et al.
Veröffentlicht: (2025)
KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation
von: Jiang, Chaoyi, et al.
Veröffentlicht: (2024)
von: Jiang, Chaoyi, et al.
Veröffentlicht: (2024)
OMPILOT: Harnessing Transformer Models for Auto Parallelization to Shared Memory Computing Paradigms
von: Bhattacharjee, Arijit, et al.
Veröffentlicht: (2025)
von: Bhattacharjee, Arijit, et al.
Veröffentlicht: (2025)
GPU Cluster Scheduling for Network-Sensitive Deep Learning
von: Sharma, Aakash, et al.
Veröffentlicht: (2024)
von: Sharma, Aakash, et al.
Veröffentlicht: (2024)
Phantora: Maximizing Code Reuse in Simulation-based Machine Learning System Performance Estimation
von: Qin, Jianxing, et al.
Veröffentlicht: (2025)
von: Qin, Jianxing, et al.
Veröffentlicht: (2025)
PyGim: An Efficient Graph Neural Network Library for Real Processing-In-Memory Architectures
von: Giannoula, Christina, et al.
Veröffentlicht: (2024)
von: Giannoula, Christina, et al.
Veröffentlicht: (2024)
SDT-GNN: Streaming-based Distributed Training Framework for Graph Neural Networks
von: Huang, Xin, et al.
Veröffentlicht: (2024)
von: Huang, Xin, et al.
Veröffentlicht: (2024)
BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures
von: Hu, Xiannan, et al.
Veröffentlicht: (2025)
von: Hu, Xiannan, et al.
Veröffentlicht: (2025)
Fake Runs, Real Fixes -- Analyzing xPU Performance Through Simulation
von: Zarkadas, Ioannis, et al.
Veröffentlicht: (2025)
von: Zarkadas, Ioannis, et al.
Veröffentlicht: (2025)
Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
von: Titopoulos, Vasileios, et al.
Veröffentlicht: (2025)
von: Titopoulos, Vasileios, et al.
Veröffentlicht: (2025)
TaxBreak: Unmasking the Hidden Costs of LLM Inference Through Overhead Decomposition
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2026)
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2026)
DeepCQ: General-Purpose Deep-Surrogate Framework for Lossy Compression Quality Prediction
von: Mumenin, Khondoker Mirazul, et al.
Veröffentlicht: (2025)
von: Mumenin, Khondoker Mirazul, et al.
Veröffentlicht: (2025)
IPA: Inference Pipeline Adaptation to Achieve High Accuracy and Cost-Efficiency
von: Ghafouri, Saeid, et al.
Veröffentlicht: (2023)
von: Ghafouri, Saeid, et al.
Veröffentlicht: (2023)
CoFormer: Collaborating with Heterogeneous Edge Devices for Scalable Transformer Inference
von: Xu, Guanyu, et al.
Veröffentlicht: (2025)
von: Xu, Guanyu, et al.
Veröffentlicht: (2025)
cedar: Optimized and Unified Machine Learning Input Data Pipelines
von: Zhao, Mark, et al.
Veröffentlicht: (2024)
von: Zhao, Mark, et al.
Veröffentlicht: (2024)
Accelerating Mobile Inference through Fine-Grained CPU-GPU Co-Execution
von: Li, Zhuojin, et al.
Veröffentlicht: (2025)
von: Li, Zhuojin, et al.
Veröffentlicht: (2025)
Multi-DNN Inference of Sparse Models on Edge SoCs
von: Luo, Jiawei, et al.
Veröffentlicht: (2026)
von: Luo, Jiawei, et al.
Veröffentlicht: (2026)
MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces
von: Sridharan, Srinivas, et al.
Veröffentlicht: (2026)
von: Sridharan, Srinivas, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Rudder: Steering Prefetching in Distributed GNN Training using LLM Agents
von: Sarkar, Aishwarya, et al.
Veröffentlicht: (2026) -
NOMAD: Generating Embeddings for Massive Distributed Graphs
von: Sarkar, Aishwarya, et al.
Veröffentlicht: (2026) -
Communication-free Sampling and 4D Hybrid Parallelism for Scalable Mini-batch GNN Training
von: Wei, Cunyang, et al.
Veröffentlicht: (2026) -
MIREncoder: Multi-modal IR-based Pretrained Embeddings for Performance Optimizations
von: Dutta, Akash, et al.
Veröffentlicht: (2024) -
InkStream: Real-time GNN Inference on Streaming Graphs via Incremental Update
von: Wu, Dan, et al.
Veröffentlicht: (2023)