Morphling: Fast, Fused, and Flexible GNN Training at Scale
Fuente:
arXiv
Saved in:
| Main Authors: | Anubhab, Nasre, Rupesh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaling Deep Learning Training with MPMD Pipeline Parallelism
by: Xhebraj, Anxhelo, et al.
Published: (2024)
by: Xhebraj, Anxhelo, et al.
Published: (2024)
StarDist: A Code Generator for Distributed Graph Algorithms
by: Nandy, Barenya Kumar, et al.
Published: (2025)
by: Nandy, Barenya Kumar, et al.
Published: (2025)
Scalable Maxflow Processing for Dynamic Graphs
by: Kannappan, Shruthi, et al.
Published: (2025)
by: Kannappan, Shruthi, et al.
Published: (2025)
veScale: Consistent and Efficient Tensor Programming with Eager-Mode SPMD
by: Li, Youjie, et al.
Published: (2025)
by: Li, Youjie, et al.
Published: (2025)
Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference
by: Yang, Mengtian, et al.
Published: (2026)
by: Yang, Mengtian, et al.
Published: (2026)
LSM-GNN: Large-scale Storage-based Multi-GPU GNN Training by Optimizing Data Transfer Scheme
by: Park, Jeongmin Brian, et al.
Published: (2024)
by: Park, Jeongmin Brian, et al.
Published: (2024)
Fused3S: Fast Sparse Attention on Tensor Cores
by: Li, Zitong, et al.
Published: (2025)
by: Li, Zitong, et al.
Published: (2025)
FastGL: A GPU-Efficient Framework for Accelerating Sampling-Based GNN Training at Large Scale
by: Zhu, Zeyu, et al.
Published: (2024)
by: Zhu, Zeyu, et al.
Published: (2024)
Theoretical Foundations of GPU-Native Compilation for Rapid Code Iteration
by: Metinov, Adilet, et al.
Published: (2025)
by: Metinov, Adilet, et al.
Published: (2025)
Mirage Persistent Kernel: A Compiler and Runtime for Mega-Kernelizing Tensor Programs
by: Cheng, Xinhao, et al.
Published: (2025)
by: Cheng, Xinhao, et al.
Published: (2025)
Data-efficient Performance Modeling via Pre-training
by: Liu, Chunting, et al.
Published: (2025)
by: Liu, Chunting, et al.
Published: (2025)
Learned Cost Model for Placement on Reconfigurable Dataflow Hardware
by: Guha, Etash, et al.
Published: (2025)
by: Guha, Etash, et al.
Published: (2025)
LOOPer: A Learned Automatic Code Optimizer For Polyhedral Compilers
by: Merouani, Massinissa, et al.
Published: (2024)
by: Merouani, Massinissa, et al.
Published: (2024)
VTC: DNN Compilation with Virtual Tensors for Data Movement Elimination
by: Hu, Muyan, et al.
Published: (2026)
by: Hu, Muyan, et al.
Published: (2026)
GPU-Accelerated Synthesis of Mixed-Boolean Arithmetic: Beyond Caching
by: Bathie, Gabriel, et al.
Published: (2026)
by: Bathie, Gabriel, et al.
Published: (2026)
PartIR: Composing SPMD Partitioning Strategies for Machine Learning
by: Alabed, Sami, et al.
Published: (2024)
by: Alabed, Sami, et al.
Published: (2024)
Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel
by: Jin, Hongyi, et al.
Published: (2026)
by: Jin, Hongyi, et al.
Published: (2026)
Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo
by: Charles, Zachary, et al.
Published: (2025)
by: Charles, Zachary, et al.
Published: (2025)
Comprehensive Evaluation of GNN Training Systems: A Data Management Perspective
by: Yuan, Hao, et al.
Published: (2023)
by: Yuan, Hao, et al.
Published: (2023)
MSPipe: Efficient Temporal GNN Training via Staleness-Aware Pipeline
by: Sheng, Guangming, et al.
Published: (2024)
by: Sheng, Guangming, et al.
Published: (2024)
Reducing Memory Contention and I/O Congestion for Disk-based GNN Training
by: Jiang, Qisheng, et al.
Published: (2024)
by: Jiang, Qisheng, et al.
Published: (2024)
SDT-GNN: Streaming-based Distributed Training Framework for Graph Neural Networks
by: Huang, Xin, et al.
Published: (2024)
by: Huang, Xin, et al.
Published: (2024)
Efficient Dynamic MaxFlow Computation on GPUs
by: Kannappan, Shruthi, et al.
Published: (2025)
by: Kannappan, Shruthi, et al.
Published: (2025)
Universal Checkpointing: A Flexible and Efficient Distributed Checkpointing System for Large-Scale DNN Training with Reconfigurable Parallelis
by: Lian, Xinyu, et al.
Published: (2024)
by: Lian, Xinyu, et al.
Published: (2024)
Agentic Auto-Scheduling: An Experimental Study of LLM-Guided Loop Optimization
by: Merouani, Massinissa, et al.
Published: (2025)
by: Merouani, Massinissa, et al.
Published: (2025)
Fast Matrix Multiplications for Lookup Table-Quantized LLMs
by: Guo, Han, et al.
Published: (2024)
by: Guo, Han, et al.
Published: (2024)
DynaTrain: Fast Online Parallelism Switching for Elastic LLM Training
by: Wang, Yuanqing, et al.
Published: (2026)
by: Wang, Yuanqing, et al.
Published: (2026)
Single-GPU GNN Systems: Traps and Pitfalls
by: Gong, Yidong, et al.
Published: (2024)
by: Gong, Yidong, et al.
Published: (2024)
MassiveGNN: Efficient Training via Prefetching for Massively Connected Distributed Graphs
by: Sarkar, Aishwarya, et al.
Published: (2024)
by: Sarkar, Aishwarya, et al.
Published: (2024)
Generating Dynamic Graph Algorithms for Multiple Backends for a Graph DSL
by: Behera, Nibedita, et al.
Published: (2025)
by: Behera, Nibedita, et al.
Published: (2025)
Scalable Training of Mixture-of-Experts Models with Megatron Core
by: Yan, Zijie, et al.
Published: (2026)
by: Yan, Zijie, et al.
Published: (2026)
Optimizing RLHF Training for Large Language Models with Stage Fusion
by: Zhong, Yinmin, et al.
Published: (2024)
by: Zhong, Yinmin, et al.
Published: (2024)
Unlocking Full Efficiency of Token Filtering in Large Language Model Training
by: Chai, Di, et al.
Published: (2025)
by: Chai, Di, et al.
Published: (2025)
Deal: Distributed End-to-End GNN Inference for All Nodes
by: Chen, Shiyang, et al.
Published: (2025)
by: Chen, Shiyang, et al.
Published: (2025)
GNNBENCH: Fair and Productive Benchmarking for Single-GPU GNN System
by: Gong, Yidong, et al.
Published: (2024)
by: Gong, Yidong, et al.
Published: (2024)
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
by: Li, Wenxuan, et al.
Published: (2025)
by: Li, Wenxuan, et al.
Published: (2025)
Echo: Simulating Distributed Training At Scale
by: Feng, Yicheng, et al.
Published: (2024)
by: Feng, Yicheng, et al.
Published: (2024)
Towards a Flexible and High-Fidelity Approach to Distributed DNN Training Emulation
by: Liu, Banruo, et al.
Published: (2024)
by: Liu, Banruo, et al.
Published: (2024)
P/D-Serve: Serving Disaggregated Large Language Model at Scale
by: Jin, Yibo, et al.
Published: (2024)
by: Jin, Yibo, et al.
Published: (2024)
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training
by: Fernandez, Jared, et al.
Published: (2024)
by: Fernandez, Jared, et al.
Published: (2024)
Similar Items
-
Scaling Deep Learning Training with MPMD Pipeline Parallelism
by: Xhebraj, Anxhelo, et al.
Published: (2024) -
StarDist: A Code Generator for Distributed Graph Algorithms
by: Nandy, Barenya Kumar, et al.
Published: (2025) -
Scalable Maxflow Processing for Dynamic Graphs
by: Kannappan, Shruthi, et al.
Published: (2025) -
veScale: Consistent and Efficient Tensor Programming with Eager-Mode SPMD
by: Li, Youjie, et al.
Published: (2025) -
Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference
by: Yang, Mengtian, et al.
Published: (2026)