NNTile: a machine learning framework capable of training extremely large GPT language models on a single node

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mikhalev, Aleksandr, Katrutsa, Aleksandr, Sozykin, Konstantin, Oseledets, Ivan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910913903198208
author Mikhalev, Aleksandr
Katrutsa, Aleksandr
Sozykin, Konstantin
Oseledets, Ivan
author_facet Mikhalev, Aleksandr
Katrutsa, Aleksandr
Sozykin, Konstantin
Oseledets, Ivan
contents This study presents an NNTile framework for training large deep neural networks in heterogeneous clusters. The NNTile is based on a StarPU library, which implements task-based parallelism and schedules all provided tasks onto all available processing units (CPUs and GPUs). It means that a particular operation, necessary to train a large neural network, can be performed on any of the CPU cores or GPU devices, depending on automatic scheduling decisions. Such an approach shifts the burden of deciding where to compute and when to communicate from a human being to an automatic decision maker, whether a simple greedy heuristic or a complex AI-based software. The performance of the presented tool for training large language models is demonstrated in extensive numerical experiments.
format Preprint
id arxiv_https___arxiv_org_abs_2504_13236
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NNTile: a machine learning framework capable of training extremely large GPT language models on a single node
Mikhalev, Aleksandr
Katrutsa, Aleksandr
Sozykin, Konstantin
Oseledets, Ivan
Machine Learning
Mathematical Software
This study presents an NNTile framework for training large deep neural networks in heterogeneous clusters. The NNTile is based on a StarPU library, which implements task-based parallelism and schedules all provided tasks onto all available processing units (CPUs and GPUs). It means that a particular operation, necessary to train a large neural network, can be performed on any of the CPU cores or GPU devices, depending on automatic scheduling decisions. Such an approach shifts the burden of deciding where to compute and when to communicate from a human being to an automatic decision maker, whether a simple greedy heuristic or a complex AI-based software. The performance of the presented tool for training large language models is demonstrated in extensive numerical experiments.
title NNTile: a machine learning framework capable of training extremely large GPT language models on a single node
topic Machine Learning
Mathematical Software
url https://arxiv.org/abs/2504.13236