Lumos: Efficient Performance Modeling and Estimation for Large-scale LLM Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Mingyu, Kassa, Hiwot Tadese, Fu, Wenyin, Coutinho, Brian, Feng, Louis, Delimitrou, Christina
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910910593892352
author Liang, Mingyu
Kassa, Hiwot Tadese
Fu, Wenyin
Coutinho, Brian
Feng, Louis
Delimitrou, Christina
author_facet Liang, Mingyu
Kassa, Hiwot Tadese
Fu, Wenyin
Coutinho, Brian
Feng, Louis
Delimitrou, Christina
contents Training LLMs in distributed environments presents significant challenges due to the complexity of model execution, deployment systems, and the vast space of configurable strategies. Although various optimization techniques exist, achieving high efficiency in practice remains difficult. Accurate performance models that effectively characterize and predict a model's behavior are essential for guiding optimization efforts and system-level studies. We propose Lumos, a trace-driven performance modeling and estimation toolkit for large-scale LLM training, designed to accurately capture and predict the execution behaviors of modern LLMs. We evaluate Lumos on a production ML cluster with up to 512 NVIDIA H100 GPUs using various GPT-3 variants, demonstrating that it can replay execution time with an average error of just 3.3%, along with other runtime details, across different models and configurations. Additionally, we validate its ability to estimate performance for new setups from existing traces, facilitating efficient exploration of model and deployment configurations.
format Preprint
id arxiv_https___arxiv_org_abs_2504_09307
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Lumos: Efficient Performance Modeling and Estimation for Large-scale LLM Training
Liang, Mingyu
Kassa, Hiwot Tadese
Fu, Wenyin
Coutinho, Brian
Feng, Louis
Delimitrou, Christina
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Training LLMs in distributed environments presents significant challenges due to the complexity of model execution, deployment systems, and the vast space of configurable strategies. Although various optimization techniques exist, achieving high efficiency in practice remains difficult. Accurate performance models that effectively characterize and predict a model's behavior are essential for guiding optimization efforts and system-level studies. We propose Lumos, a trace-driven performance modeling and estimation toolkit for large-scale LLM training, designed to accurately capture and predict the execution behaviors of modern LLMs. We evaluate Lumos on a production ML cluster with up to 512 NVIDIA H100 GPUs using various GPT-3 variants, demonstrating that it can replay execution time with an average error of just 3.3%, along with other runtime details, across different models and configurations. Additionally, we validate its ability to estimate performance for new setups from existing traces, facilitating efficient exploration of model and deployment configurations.
title Lumos: Efficient Performance Modeling and Estimation for Large-scale LLM Training
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
url https://arxiv.org/abs/2504.09307