Saved in:
| Main Authors: | Gong, Yidong, Kumar, Pradeep |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2404.04118 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Single-GPU GNN Systems: Traps and Pitfalls
by: Gong, Yidong, et al.
Published: (2024)
by: Gong, Yidong, et al.
Published: (2024)
Understanding and Reducing Metadata-Driven Host Overheads in Sampling-Based GNN Training
by: Gong, Yidong, et al.
Published: (2026)
by: Gong, Yidong, et al.
Published: (2026)
LSM-GNN: Large-scale Storage-based Multi-GPU GNN Training by Optimizing Data Transfer Scheme
by: Park, Jeongmin Brian, et al.
Published: (2024)
by: Park, Jeongmin Brian, et al.
Published: (2024)
OMEGA: A Low-Latency GNN Serving System for Large Graphs
by: Kim, Geon-Woo, et al.
Published: (2025)
by: Kim, Geon-Woo, et al.
Published: (2025)
Comprehensive Evaluation of GNN Training Systems: A Data Management Perspective
by: Yuan, Hao, et al.
Published: (2023)
by: Yuan, Hao, et al.
Published: (2023)
FastGL: A GPU-Efficient Framework for Accelerating Sampling-Based GNN Training at Large Scale
by: Zhu, Zeyu, et al.
Published: (2024)
by: Zhu, Zeyu, et al.
Published: (2024)
MoE-Gen: High-Throughput MoE Inference on a Single GPU with Module-Based Batching
by: Xu, Tairan, et al.
Published: (2025)
by: Xu, Tairan, et al.
Published: (2025)
Deal: Distributed End-to-End GNN Inference for All Nodes
by: Chen, Shiyang, et al.
Published: (2025)
by: Chen, Shiyang, et al.
Published: (2025)
Eliminating Multi-GPU Performance Taxes: A Systems Approach to Efficient Distributed LLMs
by: Trifan, Octavian Alexandru, et al.
Published: (2025)
by: Trifan, Octavian Alexandru, et al.
Published: (2025)
MSPipe: Efficient Temporal GNN Training via Staleness-Aware Pipeline
by: Sheng, Guangming, et al.
Published: (2024)
by: Sheng, Guangming, et al.
Published: (2024)
SURGE: SuperBatch Unified Resource-efficient GPU Encoding for Heterogeneous Partitioned Data
by: Kapadia, Shashank, et al.
Published: (2026)
by: Kapadia, Shashank, et al.
Published: (2026)
Revati: Transparent GPU-Free Time-Warp Emulation for LLM Serving
by: Agrawal, Amey, et al.
Published: (2026)
by: Agrawal, Amey, et al.
Published: (2026)
Reducing Memory Contention and I/O Congestion for Disk-based GNN Training
by: Jiang, Qisheng, et al.
Published: (2024)
by: Jiang, Qisheng, et al.
Published: (2024)
SDT-GNN: Streaming-based Distributed Training Framework for Graph Neural Networks
by: Huang, Xin, et al.
Published: (2024)
by: Huang, Xin, et al.
Published: (2024)
Fair Decentralized Learning
by: Biswas, Sayan, et al.
Published: (2024)
by: Biswas, Sayan, et al.
Published: (2024)
LoGoFair: Post-Processing for Local and Global Fairness in Federated Learning
by: Zhang, Li, et al.
Published: (2025)
by: Zhang, Li, et al.
Published: (2025)
FairGFL: Privacy-Preserving Fairness-Aware Federated Learning with Overlapping Subgraphs
by: Zhou, Zihao, et al.
Published: (2025)
by: Zhou, Zihao, et al.
Published: (2025)
Fair Distributed Cooperative Bandit Learning on Networks for Intelligent Internet of Things Systems (Technical Report)
by: Chen, Ziqun, et al.
Published: (2024)
by: Chen, Ziqun, et al.
Published: (2024)
KernelFoundry: Hardware-aware evolutionary GPU kernel optimization
by: Wiedemann, Nina, et al.
Published: (2026)
by: Wiedemann, Nina, et al.
Published: (2026)
Semantic-Aware Scheduling for GPU Clusters with Large Language Models
by: Wang, Zerui, et al.
Published: (2025)
by: Wang, Zerui, et al.
Published: (2025)
Iris: First-Class Multi-GPU Programming Experience in Triton
by: Awad, Muhammad, et al.
Published: (2025)
by: Awad, Muhammad, et al.
Published: (2025)
Efficient Remote KV Cache Reuse with GPU-native Video Codec
by: Mi, Liang, et al.
Published: (2026)
by: Mi, Liang, et al.
Published: (2026)
Keras Sig: Efficient Path Signature Computation on GPU in Keras 3
by: Genet, Rémi, et al.
Published: (2025)
by: Genet, Rémi, et al.
Published: (2025)
ParallelKittens: Systematic and Practical Simplification of Multi-GPU AI Kernels
by: Sul, Stuart H., et al.
Published: (2025)
by: Sul, Stuart H., et al.
Published: (2025)
Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving
by: Shi, Xiaoxiang, et al.
Published: (2025)
by: Shi, Xiaoxiang, et al.
Published: (2025)
Ilargi: a GPU Compatible Factorized ML Model Training Framework
by: Sun, Wenbo, et al.
Published: (2025)
by: Sun, Wenbo, et al.
Published: (2025)
Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference
by: Zhang, Haolin, et al.
Published: (2025)
by: Zhang, Haolin, et al.
Published: (2025)
Parallel Track Transformers: Enabling Fast GPU Inference with Reduced Synchronization
by: Wang, Chong, et al.
Published: (2026)
by: Wang, Chong, et al.
Published: (2026)
A Distributed Framework for Causal Modeling of Performance Variability in GPU Traces
by: Lahiry, Ankur, et al.
Published: (2025)
by: Lahiry, Ankur, et al.
Published: (2025)
A Unified CPU-GPU Protocol for GNN Training
by: Lin, Yi-Chien, et al.
Published: (2024)
by: Lin, Yi-Chien, et al.
Published: (2024)
Enhancing Convergence, Privacy and Fairness for Wireless Personalized Federated Learning: Quantization-Assisted Min-Max Fair Scheduling
by: Zhao, Xiyu, et al.
Published: (2025)
by: Zhao, Xiyu, et al.
Published: (2025)
Mélange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
by: Griggs, Tyler, et al.
Published: (2024)
by: Griggs, Tyler, et al.
Published: (2024)
RAPID-Serve: Resource-efficient and Accelerated P/D Intra-GPU Disaggregation
by: Masood, Amna, et al.
Published: (2026)
by: Masood, Amna, et al.
Published: (2026)
FloatSOM: GPU-Accelerated, Distributed, Topology-Flexible Self-Organizing Maps
by: Xu, Tony, et al.
Published: (2026)
by: Xu, Tony, et al.
Published: (2026)
WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving
by: Lou, Chiheng, et al.
Published: (2025)
by: Lou, Chiheng, et al.
Published: (2025)
Prediction-Assisted Online Distributed Deep Learning Workload Scheduling in GPU Clusters
by: Luo, Ziyue, et al.
Published: (2025)
by: Luo, Ziyue, et al.
Published: (2025)
Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference
by: Recasens, Pol G., et al.
Published: (2025)
by: Recasens, Pol G., et al.
Published: (2025)
Maya: Optimizing Deep Learning Training Workloads using GPU Runtime Emulation
by: Yarlagadda, Srihas, et al.
Published: (2025)
by: Yarlagadda, Srihas, et al.
Published: (2025)
GPU-Accelerated Optimization of Transformer-Based Neural Networks for Real-Time Inference
by: Mukherjee, Soutrik, et al.
Published: (2026)
by: Mukherjee, Soutrik, et al.
Published: (2026)
Accelerating Sampling and Aggregation Operations in GNN Frameworks with GPU Initiated Direct Storage Accesses
by: Park, Jeongmin Brian, et al.
Published: (2023)
by: Park, Jeongmin Brian, et al.
Published: (2023)
Similar Items
-
Single-GPU GNN Systems: Traps and Pitfalls
by: Gong, Yidong, et al.
Published: (2024) -
Understanding and Reducing Metadata-Driven Host Overheads in Sampling-Based GNN Training
by: Gong, Yidong, et al.
Published: (2026) -
LSM-GNN: Large-scale Storage-based Multi-GPU GNN Training by Optimizing Data Transfer Scheme
by: Park, Jeongmin Brian, et al.
Published: (2024) -
OMEGA: A Low-Latency GNN Serving System for Large Graphs
by: Kim, Geon-Woo, et al.
Published: (2025) -
Comprehensive Evaluation of GNN Training Systems: A Data Management Perspective
by: Yuan, Hao, et al.
Published: (2023)