Modalities, a PyTorch-native Framework For Large-scale LLM Training and Research

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lübbering, Max, Ruland, Timm, Rutmann, Richard, Stollenwerk, Felix, Fitzek, David, Fromm, Michael, Weber, Alexander, Sifa, Rafet, Flores-Herr, Nicolas, Köhler, Joachim, Ali, Mehdi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917259751981056
author Lübbering, Max
Ruland, Timm
Rutmann, Richard
Stollenwerk, Felix
Fitzek, David
Fromm, Michael
Weber, Alexander
Sifa, Rafet
Flores-Herr, Nicolas
Köhler, Joachim
Ali, Mehdi
author_facet Lübbering, Max
Ruland, Timm
Rutmann, Richard
Stollenwerk, Felix
Fitzek, David
Fromm, Michael
Weber, Alexander
Sifa, Rafet
Flores-Herr, Nicolas
Köhler, Joachim
Ali, Mehdi
contents Today's LLM (pre-) training and research workflows typically allocate a significant amount of compute to large-scale ablation studies. Despite the substantial compute costs of these ablations, existing open-source frameworks provide limited tooling for these experiments, often forcing researchers to write their own wrappers and scripts. We propose Modalities, an end-to-end PyTorch-native framework that integrates data-driven LLM research with large-scale model training from two angles. Firstly, by integrating state-of-the-art parallelization strategies, it enables both efficient pretraining and systematic ablations at trillion-token and billion-parameter scale. Secondly, Modalities adopts modular design with declarative, self-contained configuration, enabling reproducibility and extensibility levels that are difficult to achieve out-of-the-box with existing LLM training frameworks.
format Preprint
id arxiv_https___arxiv_org_abs_2602_08387
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Modalities, a PyTorch-native Framework For Large-scale LLM Training and Research
Lübbering, Max
Ruland, Timm
Rutmann, Richard
Stollenwerk, Felix
Fitzek, David
Fromm, Michael
Weber, Alexander
Sifa, Rafet
Flores-Herr, Nicolas
Köhler, Joachim
Ali, Mehdi
Machine Learning
Distributed, Parallel, and Cluster Computing
Today's LLM (pre-) training and research workflows typically allocate a significant amount of compute to large-scale ablation studies. Despite the substantial compute costs of these ablations, existing open-source frameworks provide limited tooling for these experiments, often forcing researchers to write their own wrappers and scripts. We propose Modalities, an end-to-end PyTorch-native framework that integrates data-driven LLM research with large-scale model training from two angles. Firstly, by integrating state-of-the-art parallelization strategies, it enables both efficient pretraining and systematic ablations at trillion-token and billion-parameter scale. Secondly, Modalities adopts modular design with declarative, self-contained configuration, enabling reproducibility and extensibility levels that are difficult to achieve out-of-the-box with existing LLM training frameworks.
title Modalities, a PyTorch-native Framework For Large-scale LLM Training and Research
topic Machine Learning
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2602.08387