Flare: Anomaly Diagnostics for Divergent LLM Training in GPU Clusters of Thousand-Plus Scale

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cui, Weihao, Zhang, Ji, Zhao, Han, Liu, Chao, Sha, Jian, He, Bingsheng, Guo, Minyi, Chen, Quan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908821133197312
author Cui, Weihao
Zhang, Ji
Zhao, Han
Liu, Chao
Sha, Jian
He, Bingsheng
Guo, Minyi
Chen, Quan
author_facet Cui, Weihao
Zhang, Ji
Zhao, Han
Liu, Chao
Sha, Jian
He, Bingsheng
Guo, Minyi
Chen, Quan
contents The rapid proliferation of large language models has driven the need for efficient GPU training clusters. However, it is challenging due to the frequent occurrence of training anomalies. Since existing diagnostic tools are narrowly tailored to specific issues, there are gaps in their ability to address anomalies spanning the entire training stack. In response, we introduce Flare, a diagnostic framework designed for distributed LLM training at scale. Flare first integrates a lightweight tracing daemon for full-stack and backend-extensible tracing. Additionally, it features a diagnostic engine that automatically diagnoses anomalies, with a focus on performance regressions. The deployment of Flare across 6,000 GPUs has demonstrated significant improvements in pinpointing deficiencies in real-world scenarios, with continuous operation for over eight months.
format Preprint
id arxiv_https___arxiv_org_abs_2502_05413
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Flare: Anomaly Diagnostics for Divergent LLM Training in GPU Clusters of Thousand-Plus Scale
Cui, Weihao
Zhang, Ji
Zhao, Han
Liu, Chao
Sha, Jian
He, Bingsheng
Guo, Minyi
Chen, Quan
Operating Systems
The rapid proliferation of large language models has driven the need for efficient GPU training clusters. However, it is challenging due to the frequent occurrence of training anomalies. Since existing diagnostic tools are narrowly tailored to specific issues, there are gaps in their ability to address anomalies spanning the entire training stack. In response, we introduce Flare, a diagnostic framework designed for distributed LLM training at scale. Flare first integrates a lightweight tracing daemon for full-stack and backend-extensible tracing. Additionally, it features a diagnostic engine that automatically diagnoses anomalies, with a focus on performance regressions. The deployment of Flare across 6,000 GPUs has demonstrated significant improvements in pinpointing deficiencies in real-world scenarios, with continuous operation for over eight months.
title Flare: Anomaly Diagnostics for Divergent LLM Training in GPU Clusters of Thousand-Plus Scale
topic Operating Systems
url https://arxiv.org/abs/2502.05413