Morphlux: Transforming Torus Fabrics for Efficient Multi-tenant ML

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kumar, Abhishek Vijaya, Ding, Eric, Devraj, Arjun, Bunandar, Darius, Singh, Rachee
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915531519426560
author Kumar, Abhishek Vijaya
Ding, Eric
Devraj, Arjun
Bunandar, Darius
Singh, Rachee
author_facet Kumar, Abhishek Vijaya
Ding, Eric
Devraj, Arjun
Bunandar, Darius
Singh, Rachee
contents We develop Morphlux, a server-scale programmable photonic fabric to interconnect accelerators within servers. We show that augmenting state-of-the-art torus-based ML data-centers with Morphlux can improve the bandwidth of tenant compute allocations by up to 66%, reduce compute fragmentation by up to 70%, and minimize the blast radius of chip failures. We develop a novel end-to-end hardware prototype of Morphlux to demonstrate these performance benefits which translate to 1.72X improvement in training throughput of ML models. By rapidly programming the server-scale fabric in our hardware testbed, Morphlux can replace a failed accelerator chip with a healthy one in 1.2 seconds.
format Preprint
id arxiv_https___arxiv_org_abs_2508_03674
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Morphlux: Transforming Torus Fabrics for Efficient Multi-tenant ML
Kumar, Abhishek Vijaya
Ding, Eric
Devraj, Arjun
Bunandar, Darius
Singh, Rachee
Networking and Internet Architecture
Hardware Architecture
Machine Learning
We develop Morphlux, a server-scale programmable photonic fabric to interconnect accelerators within servers. We show that augmenting state-of-the-art torus-based ML data-centers with Morphlux can improve the bandwidth of tenant compute allocations by up to 66%, reduce compute fragmentation by up to 70%, and minimize the blast radius of chip failures. We develop a novel end-to-end hardware prototype of Morphlux to demonstrate these performance benefits which translate to 1.72X improvement in training throughput of ML models. By rapidly programming the server-scale fabric in our hardware testbed, Morphlux can replace a failed accelerator chip with a healthy one in 1.2 seconds.
title Morphlux: Transforming Torus Fabrics for Efficient Multi-tenant ML
topic Networking and Internet Architecture
Hardware Architecture
Machine Learning
url https://arxiv.org/abs/2508.03674