Accelerate Model Parallel Training by Using Efficient Graph Traversal Order in Device Placement
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Tianze, Payberah, Amir H., Hagos, Desta Haileselassie, Vlassov, Vladimir |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Recent Advances in Generative AI and Large Language Models: Current Status, Challenges, and Perspectives
by: Hagos, Desta Haileselassie, et al.
Published: (2024)
by: Hagos, Desta Haileselassie, et al.
Published: (2024)
BiSparse-AAS: Bilinear Sparse Attention and Adaptive Spans Framework for Scalable and Efficient Text Summarization
by: Hagos, Desta Haileselassie, et al.
Published: (2025)
by: Hagos, Desta Haileselassie, et al.
Published: (2025)
Neuro-Symbolic AI for Military Applications
by: Hagos, Desta Haileselassie, et al.
Published: (2024)
by: Hagos, Desta Haileselassie, et al.
Published: (2024)
An Adaptive Placement and Parallelism Framework for Accelerating RLHF Training
by: Xiao, Youshao, et al.
Published: (2023)
by: Xiao, Youshao, et al.
Published: (2023)
Metaverse Survey & Tutorial: Exploring Key Requirements, Technologies, Standards, Applications, Challenges, and Perspectives
by: Rawat, Danda B., et al.
Published: (2024)
by: Rawat, Danda B., et al.
Published: (2024)
AI-Driven Human-Autonomy Teaming in Tactical Operations: Proposed Framework, Challenges, and Future Directions
by: Hagos, Desta Haileselassie, et al.
Published: (2024)
by: Hagos, Desta Haileselassie, et al.
Published: (2024)
The Impact of Background Removal on Performance of Neural Networks for Fashion Image Classification and Segmentation
by: Liang, Junhui, et al.
Published: (2023)
by: Liang, Junhui, et al.
Published: (2023)
MixGCN: Scalable GCN Training by Mixture of Parallelism and Mixture of Accelerators
by: Wan, Cheng, et al.
Published: (2025)
by: Wan, Cheng, et al.
Published: (2025)
Efficient Zero-Order Federated Finetuning of Language Models for Resource-Constrained Devices
by: Ahmed, Mohamed Aboelenien, et al.
Published: (2025)
by: Ahmed, Mohamed Aboelenien, et al.
Published: (2025)
Chiplet Placement Order Exploration Based on Learning to Rank with Graph Representation
by: Deng, Zhihui, et al.
Published: (2024)
by: Deng, Zhihui, et al.
Published: (2024)
Fast Graph Generation via Spectral Diffusion
by: Luo, Tianze, et al.
Published: (2022)
by: Luo, Tianze, et al.
Published: (2022)
ASTRA: Communication-Efficient Acceleration for Multi-Device Transformer Inference
by: Liu, Xiao, et al.
Published: (2025)
by: Liu, Xiao, et al.
Published: (2025)
MoNTA: Accelerating Mixture-of-Experts Training with Network-Traffc-Aware Parallel Optimization
by: Guo, Jingming, et al.
Published: (2024)
by: Guo, Jingming, et al.
Published: (2024)
Graph Counterfactual Explainable AI via Latent Space Traversal
by: Hansen, Andreas Abildtrup, et al.
Published: (2025)
by: Hansen, Andreas Abildtrup, et al.
Published: (2025)
Accelerating Storage-Based Training for Graph Neural Networks
by: Jang, Myung-Hwan, et al.
Published: (2026)
by: Jang, Myung-Hwan, et al.
Published: (2026)
SlimPipe: Memory-Thrifty and Efficient Pipeline Parallelism for Long-Context LLM Training
by: Li, Zhouyang, et al.
Published: (2025)
by: Li, Zhouyang, et al.
Published: (2025)
LGAN: An Efficient High-Order Graph Neural Network via the Line Graph Aggregation
by: Du, Lin, et al.
Published: (2025)
by: Du, Lin, et al.
Published: (2025)
Task-Oriented GNNs Training on Large Knowledge Graphs for Accurate and Efficient Modeling
by: Abdallah, Hussein, et al.
Published: (2024)
by: Abdallah, Hussein, et al.
Published: (2024)
Rethinking and Accelerating Graph Condensation: A Training-Free Approach with Class Partition
by: Gao, Xinyi, et al.
Published: (2024)
by: Gao, Xinyi, et al.
Published: (2024)
Understanding and Guiding Layer Placement in Parameter-Efficient Fine-Tuning of Large Language Models
by: Xu, Yichen, et al.
Published: (2026)
by: Xu, Yichen, et al.
Published: (2026)
BitPipe: Bidirectional Interleaved Pipeline Parallelism for Accelerating Large Models Training
by: Wu, Houming, et al.
Published: (2024)
by: Wu, Houming, et al.
Published: (2024)
GraphiT: Efficient Node Classification on Text-Attributed Graphs with Prompt Optimized LLMs
by: Khoshraftar, Shima, et al.
Published: (2025)
by: Khoshraftar, Shima, et al.
Published: (2025)
Efficient Knowledge Tracing Leveraging Higher-Order Information in Integrated Graphs
by: Han, Donghee, et al.
Published: (2025)
by: Han, Donghee, et al.
Published: (2025)
GSR-GNN: Training Acceleration and Memory-Saving Framework of Deep GNNs on Circuit Graph
by: Luo, Yuebo, et al.
Published: (2026)
by: Luo, Yuebo, et al.
Published: (2026)
Energy Consumption in Parallel Neural Network Training
by: Huber, Philipp, et al.
Published: (2025)
by: Huber, Philipp, et al.
Published: (2025)
Pgx: Hardware-Accelerated Parallel Game Simulators for Reinforcement Learning
by: Koyamada, Sotetsu, et al.
Published: (2023)
by: Koyamada, Sotetsu, et al.
Published: (2023)
Scalable and Adaptive Parallel Training of Graph Transformer on Large Graphs
by: Lin, Jun-Liang, et al.
Published: (2026)
by: Lin, Jun-Liang, et al.
Published: (2026)
GraphProp: Training the Graph Foundation Models using Graph Properties
by: Sun, Ziheng, et al.
Published: (2025)
by: Sun, Ziheng, et al.
Published: (2025)
Advancing On-Device Neural Network Training with TinyPropv2: Dynamic, Sparse, and Efficient Backpropagation
by: Rüb, Marcus, et al.
Published: (2024)
by: Rüb, Marcus, et al.
Published: (2024)
Model Parallelism With Subnetwork Data Parallelism
by: Singh, Vaibhav, et al.
Published: (2025)
by: Singh, Vaibhav, et al.
Published: (2025)
OptEx: Expediting First-Order Optimization with Approximately Parallelized Iterations
by: Shu, Yao, et al.
Published: (2024)
by: Shu, Yao, et al.
Published: (2024)
Lightweight and Post-Training Structured Pruning for On-Device Large Lanaguage Models
by: Xu, Zihuai, et al.
Published: (2025)
by: Xu, Zihuai, et al.
Published: (2025)
Athena: Efficient Block-Wise Post-Training Quantization for Large Language Models Using Second-Order Matrix Derivative Information
by: Wang, Yanshu, et al.
Published: (2024)
by: Wang, Yanshu, et al.
Published: (2024)
SCPL: Enhancing Neural Network Training Throughput with Decoupled Local Losses and Model Parallelism
by: Ho, Ming-Yao, et al.
Published: (2026)
by: Ho, Ming-Yao, et al.
Published: (2026)
Bifurcated Attention: Accelerating Massively Parallel Decoding with Shared Prefixes in LLMs
by: Athiwaratkun, Ben, et al.
Published: (2024)
by: Athiwaratkun, Ben, et al.
Published: (2024)
P-EAGLE: Parallel-Drafting EAGLE with Scalable Training
by: Hui, Mude, et al.
Published: (2026)
by: Hui, Mude, et al.
Published: (2026)
Learning-Order Autoregressive Models with Application to Molecular Graph Generation
by: Wang, Zhe, et al.
Published: (2025)
by: Wang, Zhe, et al.
Published: (2025)
Decentralized Adversarial Training over Graphs
by: Cao, Ying, et al.
Published: (2023)
by: Cao, Ying, et al.
Published: (2023)
Efficient Causal Graph Discovery Using Large Language Models
by: Jiralerspong, Thomas, et al.
Published: (2024)
by: Jiralerspong, Thomas, et al.
Published: (2024)
Leveraging Machine Learning Models to Predict the Outcome of Digital Medical Triage Interviews
by: Krylova, Sofia, et al.
Published: (2025)
by: Krylova, Sofia, et al.
Published: (2025)
Similar Items
-
Recent Advances in Generative AI and Large Language Models: Current Status, Challenges, and Perspectives
by: Hagos, Desta Haileselassie, et al.
Published: (2024) -
BiSparse-AAS: Bilinear Sparse Attention and Adaptive Spans Framework for Scalable and Efficient Text Summarization
by: Hagos, Desta Haileselassie, et al.
Published: (2025) -
Neuro-Symbolic AI for Military Applications
by: Hagos, Desta Haileselassie, et al.
Published: (2024) -
An Adaptive Placement and Parallelism Framework for Accelerating RLHF Training
by: Xiao, Youshao, et al.
Published: (2023) -
Metaverse Survey & Tutorial: Exploring Key Requirements, Technologies, Standards, Applications, Challenges, and Perspectives
by: Rawat, Danda B., et al.
Published: (2024)