Characterizing and Understanding HGNN Training on GPUs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Han, Dengke, Yan, Mingyu, Ye, Xiaochun, Fan, Dongrui
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909369912786944
author Han, Dengke
Yan, Mingyu
Ye, Xiaochun
Fan, Dongrui
author_facet Han, Dengke
Yan, Mingyu
Ye, Xiaochun
Fan, Dongrui
contents Owing to their remarkable representation capabilities for heterogeneous graph data, Heterogeneous Graph Neural Networks (HGNNs) have been widely adopted in many critical real-world domains such as recommendation systems and medical analysis. Prior to their practical application, identifying the optimal HGNN model parameters tailored to specific tasks through extensive training is a time-consuming and costly process. To enhance the efficiency of HGNN training, it is essential to characterize and analyze the execution semantics and patterns within the training process to identify performance bottlenecks. In this study, we conduct an in-depth quantification and analysis of two mainstream HGNN training scenarios, including single-GPU and multi-GPU distributed training. Based on the characterization results, we disclose the performance bottlenecks and their underlying causes in different HGNN training scenarios and provide optimization guidelines from both software and hardware perspectives.
format Preprint
id arxiv_https___arxiv_org_abs_2407_11790
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Characterizing and Understanding HGNN Training on GPUs
Han, Dengke
Yan, Mingyu
Ye, Xiaochun
Fan, Dongrui
Machine Learning
Artificial Intelligence
Hardware Architecture
Performance
Owing to their remarkable representation capabilities for heterogeneous graph data, Heterogeneous Graph Neural Networks (HGNNs) have been widely adopted in many critical real-world domains such as recommendation systems and medical analysis. Prior to their practical application, identifying the optimal HGNN model parameters tailored to specific tasks through extensive training is a time-consuming and costly process. To enhance the efficiency of HGNN training, it is essential to characterize and analyze the execution semantics and patterns within the training process to identify performance bottlenecks. In this study, we conduct an in-depth quantification and analysis of two mainstream HGNN training scenarios, including single-GPU and multi-GPU distributed training. Based on the characterization results, we disclose the performance bottlenecks and their underlying causes in different HGNN training scenarios and provide optimization guidelines from both software and hardware perspectives.
title Characterizing and Understanding HGNN Training on GPUs
topic Machine Learning
Artificial Intelligence
Hardware Architecture
Performance
url https://arxiv.org/abs/2407.11790