LUMION: Fast Fault Recovery for ML Jobs Using Programmable Optical Fabrics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kumar, Abhishek Vijaya, Ding, Eric, Devraj, Arjun, Bunandar, Darius, Singh, Rachee
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909825056636928
author Kumar, Abhishek Vijaya
Ding, Eric
Devraj, Arjun
Bunandar, Darius
Singh, Rachee
author_facet Kumar, Abhishek Vijaya
Ding, Eric
Devraj, Arjun
Bunandar, Darius
Singh, Rachee
contents When accelerators fail in modern ML datacenters, operators migrate the affected ML training or inference jobs to entirely new racks. This approach, while preserving network performance, is highly inefficient, requiring datacenters to reserve full racks of idle accelerators for fault tolerance. In this paper, we address this resource inefficiency by introducing LUMION, a novel reconfigurable optical fabric for connecting accelerators within a datacenter rack. Instead of migrating entire ML jobs, LUMION dynamically integrates spare accelerators into ongoing workloads as failures occur, thereby maintaining consistent performance without costly migrations. We show the benefits of LUMION by building an end-to-end hardware prototype. Our experiments fine-tune Llama 3.2 and show that LUMION swaps a failed GPU with a healthy one and restarts the ML job within ~ 1 second of the failure. LUMION achieves higher inter-GPU bandwidth compared to traditional electrical racks after replacing failed accelerators with spare ones, leading to nearly 2X improvement in fine-tuning throughput.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23105
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LUMION: Fast Fault Recovery for ML Jobs Using Programmable Optical Fabrics
Kumar, Abhishek Vijaya
Ding, Eric
Devraj, Arjun
Bunandar, Darius
Singh, Rachee
Machine Learning
Networking and Internet Architecture
When accelerators fail in modern ML datacenters, operators migrate the affected ML training or inference jobs to entirely new racks. This approach, while preserving network performance, is highly inefficient, requiring datacenters to reserve full racks of idle accelerators for fault tolerance. In this paper, we address this resource inefficiency by introducing LUMION, a novel reconfigurable optical fabric for connecting accelerators within a datacenter rack. Instead of migrating entire ML jobs, LUMION dynamically integrates spare accelerators into ongoing workloads as failures occur, thereby maintaining consistent performance without costly migrations. We show the benefits of LUMION by building an end-to-end hardware prototype. Our experiments fine-tune Llama 3.2 and show that LUMION swaps a failed GPU with a healthy one and restarts the ML job within ~ 1 second of the failure. LUMION achieves higher inter-GPU bandwidth compared to traditional electrical racks after replacing failed accelerators with spare ones, leading to nearly 2X improvement in fine-tuning throughput.
title LUMION: Fast Fault Recovery for ML Jobs Using Programmable Optical Fabrics
topic Machine Learning
Networking and Internet Architecture
url https://arxiv.org/abs/2505.23105