Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Penke, Carolin, John, Chelsea Maria, Ebert, Jan, Kesselheim, Stefan, Herten, Andreas
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918487213998080
author Penke, Carolin
John, Chelsea Maria
Ebert, Jan
Kesselheim, Stefan
Herten, Andreas
author_facet Penke, Carolin
John, Chelsea Maria
Ebert, Jan
Kesselheim, Stefan
Herten, Andreas
contents The training of large language models (LLMs) requires substantial computational resources, complex software stacks, and carefully designed workflows to achieve scalability and efficiency. This report presents best practices and insights gained from the OpenGPT-X project, a German initiative focused on developing open, multilingual LLMs optimized for European languages. We detail the use of high-performance computing (HPC) systems, primarily JUWELS Booster at JSC, for training Teuken-7B, a 7-billion-parameter transformer model. The report covers system architecture, training infrastructure, software choices, profiling and benchmarking tools, as well as engineering and operational challenges. It includes measured throughput data of various configurations of 3D parallelism during training and the impact of features such as flash attention.
format Preprint
id arxiv_https___arxiv_org_abs_2504_10013
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
Penke, Carolin
John, Chelsea Maria
Ebert, Jan
Kesselheim, Stefan
Herten, Andreas
Distributed, Parallel, and Cluster Computing
C.4; I.2.11; I.2.7; K.6
The training of large language models (LLMs) requires substantial computational resources, complex software stacks, and carefully designed workflows to achieve scalability and efficiency. This report presents best practices and insights gained from the OpenGPT-X project, a German initiative focused on developing open, multilingual LLMs optimized for European languages. We detail the use of high-performance computing (HPC) systems, primarily JUWELS Booster at JSC, for training Teuken-7B, a 7-billion-parameter transformer model. The report covers system architecture, training infrastructure, software choices, profiling and benchmarking tools, as well as engineering and operational challenges. It includes measured throughput data of various configurations of 3D parallelism during training and the impact of features such as flash attention.
title Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
topic Distributed, Parallel, and Cluster Computing
C.4; I.2.11; I.2.7; K.6
url https://arxiv.org/abs/2504.10013