Comprehensive Evaluation of GNN Training Systems: A Data Management Perspective
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2023
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866910374519898112 |
|---|---|
| author | Yuan, Hao Liu, Yajiong Zhang, Yanfeng Ai, Xin Wang, Qiange Chen, Chaoyi Gu, Yu Yu, Ge |
| author_facet | Yuan, Hao Liu, Yajiong Zhang, Yanfeng Ai, Xin Wang, Qiange Chen, Chaoyi Gu, Yu Yu, Ge |
| contents | Many Graph Neural Network (GNN) training systems have emerged recently to support efficient GNN training. Since GNNs embody complex data dependencies between training samples, the training of GNNs should address distinct challenges different from DNN training in data management, such as data partitioning, batch preparation for mini-batch training, and data transferring between CPUs and GPUs. These factors, which take up a large proportion of training time, make data management in GNN training more significant. This paper reviews GNN training from a data management perspective and provides a comprehensive analysis and evaluation of the representative approaches. We conduct extensive experiments on various benchmark datasets and show many interesting and valuable results. We also provide some practical tips learned from these experiments, which are helpful for designing GNN training systems in the future. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2311_13279 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Comprehensive Evaluation of GNN Training Systems: A Data Management Perspective Yuan, Hao Liu, Yajiong Zhang, Yanfeng Ai, Xin Wang, Qiange Chen, Chaoyi Gu, Yu Yu, Ge Machine Learning Distributed, Parallel, and Cluster Computing Many Graph Neural Network (GNN) training systems have emerged recently to support efficient GNN training. Since GNNs embody complex data dependencies between training samples, the training of GNNs should address distinct challenges different from DNN training in data management, such as data partitioning, batch preparation for mini-batch training, and data transferring between CPUs and GPUs. These factors, which take up a large proportion of training time, make data management in GNN training more significant. This paper reviews GNN training from a data management perspective and provides a comprehensive analysis and evaluation of the representative approaches. We conduct extensive experiments on various benchmark datasets and show many interesting and valuable results. We also provide some practical tips learned from these experiments, which are helpful for designing GNN training systems in the future. |
| title | Comprehensive Evaluation of GNN Training Systems: A Data Management Perspective |
| topic | Machine Learning Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2311.13279 |