L4: Diagnosing Large-scale LLM Training Failures via Automated Log Analysis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jiang, Zhihan, Huang, Junjie, Chen, Zhuangbin, Li, Yichen, Yu, Guangba, Feng, Cong, Yang, Yongqiang, Yang, Zengyin, Lyu, Michael R. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TraceMesh: Scalable and Streaming Sampling for Distributed Traces
von: Chen, Zhuangbin, et al.
Veröffentlicht: (2024)
von: Chen, Zhuangbin, et al.
Veröffentlicht: (2024)
AI-NativeBench: An Open-Source White-Box Agentic Benchmark Suite for AI-Native Systems
von: Wang, Zirui, et al.
Veröffentlicht: (2026)
von: Wang, Zirui, et al.
Veröffentlicht: (2026)
AlertGuardian: Intelligent Alert Life-Cycle Management for Large-scale Cloud Systems
von: Yu, Guangba, et al.
Veröffentlicht: (2026)
von: Yu, Guangba, et al.
Veröffentlicht: (2026)
Why Does the LLM Stop Computing: An Empirical Study of User-Reported Failures in Open-Source LLMs
von: Yu, Guangba, et al.
Veröffentlicht: (2026)
von: Yu, Guangba, et al.
Veröffentlicht: (2026)
Hierarchical Prediction-based Management for LMaaS Systems
von: Jiang, Zhihan, et al.
Veröffentlicht: (2025)
von: Jiang, Zhihan, et al.
Veröffentlicht: (2025)
Metronome: Differentiated Delay Scheduling for Serverless Functions
von: Chen, Zhuangbin, et al.
Veröffentlicht: (2025)
von: Chen, Zhuangbin, et al.
Veröffentlicht: (2025)
A Survey on Failure Analysis and Fault Injection in AI Systems
von: Yu, Guangba, et al.
Veröffentlicht: (2024)
von: Yu, Guangba, et al.
Veröffentlicht: (2024)
CodeAD: Synthesize Code of Rules for Log-based Anomaly Detection with LLMs
von: Huang, Junjie, et al.
Veröffentlicht: (2025)
von: Huang, Junjie, et al.
Veröffentlicht: (2025)
MegaFlow: Large-Scale Distributed Orchestration System for the Agentic Era
von: Zhang, Lei, et al.
Veröffentlicht: (2026)
von: Zhang, Lei, et al.
Veröffentlicht: (2026)
SPES: Towards Optimizing Performance-Resource Trade-Off for Serverless Functions
von: Lee, Cheryl, et al.
Veröffentlicht: (2024)
von: Lee, Cheryl, et al.
Veröffentlicht: (2024)
CSnake: Detecting Self-Sustaining Cascading Failure via Causal Stitching of Fault Propagations
von: Qian, Shangshu, et al.
Veröffentlicht: (2025)
von: Qian, Shangshu, et al.
Veröffentlicht: (2025)
SWIFT: Expedited Failure Recovery for Large-scale DNN Training
von: Zhong, Yuchen, et al.
Veröffentlicht: (2023)
von: Zhong, Yuchen, et al.
Veröffentlicht: (2023)
LLOR: Automated Repair of OpenMP Programs
von: Bora, Utpal, et al.
Veröffentlicht: (2024)
von: Bora, Utpal, et al.
Veröffentlicht: (2024)
Performant Automatic BLAS Offloading on Unified Memory Architecture with OpenMP First-Touch Style Data Movement
von: Li, Junjie
Veröffentlicht: (2024)
von: Li, Junjie
Veröffentlicht: (2024)
Umbilical Choir: Automated Live Testing for Edge-To-Cloud FaaS Applications
von: Malekabbasi, Mohammadreza, et al.
Veröffentlicht: (2025)
von: Malekabbasi, Mohammadreza, et al.
Veröffentlicht: (2025)
LLM-HPC++: Evaluating LLM-Generated Modern C++ and MPI+OpenMP Codes for Scalable Mandelbrot Set Computation
von: Diehl, Patrick, et al.
Veröffentlicht: (2025)
von: Diehl, Patrick, et al.
Veröffentlicht: (2025)
A Framework for Effective Invocation Methods of Various LLM Services
von: Wang, Can, et al.
Veröffentlicht: (2024)
von: Wang, Can, et al.
Veröffentlicht: (2024)
ATOM: Asynchronous Training of Massive Models for Deep Learning in a Decentralized Environment
von: Wu, Xiaofeng, et al.
Veröffentlicht: (2024)
von: Wu, Xiaofeng, et al.
Veröffentlicht: (2024)
LLM4FaaS: No-Code Application Development using LLMs and FaaS
von: Wang, Minghe, et al.
Veröffentlicht: (2025)
von: Wang, Minghe, et al.
Veröffentlicht: (2025)
Supercharging Federated Learning with Flower and NVIDIA FLARE
von: Roth, Holger R., et al.
Veröffentlicht: (2024)
von: Roth, Holger R., et al.
Veröffentlicht: (2024)
Do Large Language Models Understand Performance Optimization?
von: Cui, Bowen, et al.
Veröffentlicht: (2025)
von: Cui, Bowen, et al.
Veröffentlicht: (2025)
GitFarm: Git as a Service for Large-Scale Monorepos
von: Dwivedi, Preetam, et al.
Veröffentlicht: (2026)
von: Dwivedi, Preetam, et al.
Veröffentlicht: (2026)
A Large-Scale Exploratory Study on the Proxy Pattern in Ethereum
von: Ebrahimi, Amir M., et al.
Veröffentlicht: (2025)
von: Ebrahimi, Amir M., et al.
Veröffentlicht: (2025)
CloudHeatMap: Heatmap-Based Monitoring for Large-Scale Cloud Systems
von: Sohana, Sarah, et al.
Veröffentlicht: (2024)
von: Sohana, Sarah, et al.
Veröffentlicht: (2024)
ShuffleBench: A Benchmark for Large-Scale Data Shuffling Operations with Distributed Stream Processing Frameworks
von: Henning, Sören, et al.
Veröffentlicht: (2024)
von: Henning, Sören, et al.
Veröffentlicht: (2024)
FalconGEMM: Surpassing Hardware Peaks with Lower-Complexity Matrix Multiplication
von: Zhu, Honglin, et al.
Veröffentlicht: (2026)
von: Zhu, Honglin, et al.
Veröffentlicht: (2026)
Automated MPI-X code generation for scalable finite-difference solvers
von: Bisbas, George, et al.
Veröffentlicht: (2023)
von: Bisbas, George, et al.
Veröffentlicht: (2023)
CloudFix: Automated Policy Repair for Cloud Access Control Policies Using Large Language Models
von: Hall, Bethel, et al.
Veröffentlicht: (2025)
von: Hall, Bethel, et al.
Veröffentlicht: (2025)
PALM: A Efficient Performance Simulator for Tiled Accelerators with Large-scale Model Training
von: Fang, Jiahao, et al.
Veröffentlicht: (2024)
von: Fang, Jiahao, et al.
Veröffentlicht: (2024)
ReviveMoE: Fast Recovery for Hardware Failures in Large-Scale MoE LLM Inference Deployments
von: Li, Haley, et al.
Veröffentlicht: (2026)
von: Li, Haley, et al.
Veröffentlicht: (2026)
ResiHP: Taming LLM Training Failures with Dynamic Hybrid Parallelism
von: Ma, Tenghui, et al.
Veröffentlicht: (2026)
von: Ma, Tenghui, et al.
Veröffentlicht: (2026)
LO2: Microservice API Anomaly Dataset of Logs and Metrics
von: Bakhtin, Alexander, et al.
Veröffentlicht: (2025)
von: Bakhtin, Alexander, et al.
Veröffentlicht: (2025)
Lumos: Efficient Performance Modeling and Estimation for Large-scale LLM Training
von: Liang, Mingyu, et al.
Veröffentlicht: (2025)
von: Liang, Mingyu, et al.
Veröffentlicht: (2025)
NApy: Efficient Statistics in Python for Large-Scale Heterogeneous Data with Enhanced Support for Missing Data
von: Woller, Fabian, et al.
Veröffentlicht: (2025)
von: Woller, Fabian, et al.
Veröffentlicht: (2025)
LogAction: Consistent Cross-system Anomaly Detection through Logs via Active Domain Adaptation
von: Duan, Chiming, et al.
Veröffentlicht: (2025)
von: Duan, Chiming, et al.
Veröffentlicht: (2025)
SeBS-Flow: Benchmarking Serverless Cloud Function Workflows
von: Schmid, Larissa, et al.
Veröffentlicht: (2024)
von: Schmid, Larissa, et al.
Veröffentlicht: (2024)
A Unifying Framework to Enable Artificial Intelligence in High Performance Computing Workflows
von: Domke, Jens, et al.
Veröffentlicht: (2025)
von: Domke, Jens, et al.
Veröffentlicht: (2025)
Integrating Odeint Time Stepping into OpenFPM for Distributed and GPU Accelerated Numerical Solvers
von: Singh, Abhinav, et al.
Veröffentlicht: (2023)
von: Singh, Abhinav, et al.
Veröffentlicht: (2023)
$μ$OpTime: Statically Reducing the Execution Time of Microbenchmark Suites Using Stability Metrics
von: Japke, Nils, et al.
Veröffentlicht: (2025)
von: Japke, Nils, et al.
Veröffentlicht: (2025)
Adaptable TeaStore
von: Bliudze, Simon, et al.
Veröffentlicht: (2024)
von: Bliudze, Simon, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
TraceMesh: Scalable and Streaming Sampling for Distributed Traces
von: Chen, Zhuangbin, et al.
Veröffentlicht: (2024) -
AI-NativeBench: An Open-Source White-Box Agentic Benchmark Suite for AI-Native Systems
von: Wang, Zirui, et al.
Veröffentlicht: (2026) -
AlertGuardian: Intelligent Alert Life-Cycle Management for Large-scale Cloud Systems
von: Yu, Guangba, et al.
Veröffentlicht: (2026) -
Why Does the LLM Stop Computing: An Empirical Study of User-Reported Failures in Open-Source LLMs
von: Yu, Guangba, et al.
Veröffentlicht: (2026) -
Hierarchical Prediction-based Management for LMaaS Systems
von: Jiang, Zhihan, et al.
Veröffentlicht: (2025)