Modalities, a PyTorch-native Framework For Large-scale LLM Training and Research
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917259751981056 |
|---|---|
| author | Lübbering, Max Ruland, Timm Rutmann, Richard Stollenwerk, Felix Fitzek, David Fromm, Michael Weber, Alexander Sifa, Rafet Flores-Herr, Nicolas Köhler, Joachim Ali, Mehdi |
| author_facet | Lübbering, Max Ruland, Timm Rutmann, Richard Stollenwerk, Felix Fitzek, David Fromm, Michael Weber, Alexander Sifa, Rafet Flores-Herr, Nicolas Köhler, Joachim Ali, Mehdi |
| contents | Today's LLM (pre-) training and research workflows typically allocate a significant amount of compute to large-scale ablation studies. Despite the substantial compute costs of these ablations, existing open-source frameworks provide limited tooling for these experiments, often forcing researchers to write their own wrappers and scripts. We propose Modalities, an end-to-end PyTorch-native framework that integrates data-driven LLM research with large-scale model training from two angles. Firstly, by integrating state-of-the-art parallelization strategies, it enables both efficient pretraining and systematic ablations at trillion-token and billion-parameter scale. Secondly, Modalities adopts modular design with declarative, self-contained configuration, enabling reproducibility and extensibility levels that are difficult to achieve out-of-the-box with existing LLM training frameworks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_08387 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Modalities, a PyTorch-native Framework For Large-scale LLM Training and Research Lübbering, Max Ruland, Timm Rutmann, Richard Stollenwerk, Felix Fitzek, David Fromm, Michael Weber, Alexander Sifa, Rafet Flores-Herr, Nicolas Köhler, Joachim Ali, Mehdi Machine Learning Distributed, Parallel, and Cluster Computing Today's LLM (pre-) training and research workflows typically allocate a significant amount of compute to large-scale ablation studies. Despite the substantial compute costs of these ablations, existing open-source frameworks provide limited tooling for these experiments, often forcing researchers to write their own wrappers and scripts. We propose Modalities, an end-to-end PyTorch-native framework that integrates data-driven LLM research with large-scale model training from two angles. Firstly, by integrating state-of-the-art parallelization strategies, it enables both efficient pretraining and systematic ablations at trillion-token and billion-parameter scale. Secondly, Modalities adopts modular design with declarative, self-contained configuration, enabling reproducibility and extensibility levels that are difficult to achieve out-of-the-box with existing LLM training frameworks. |
| title | Modalities, a PyTorch-native Framework For Large-scale LLM Training and Research |
| topic | Machine Learning Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2602.08387 |