Morphlux: Transforming Torus Fabrics for Efficient Multi-tenant ML
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915531519426560 |
|---|---|
| author | Kumar, Abhishek Vijaya Ding, Eric Devraj, Arjun Bunandar, Darius Singh, Rachee |
| author_facet | Kumar, Abhishek Vijaya Ding, Eric Devraj, Arjun Bunandar, Darius Singh, Rachee |
| contents | We develop Morphlux, a server-scale programmable photonic fabric to interconnect accelerators within servers. We show that augmenting state-of-the-art torus-based ML data-centers with Morphlux can improve the bandwidth of tenant compute allocations by up to 66%, reduce compute fragmentation by up to 70%, and minimize the blast radius of chip failures. We develop a novel end-to-end hardware prototype of Morphlux to demonstrate these performance benefits which translate to 1.72X improvement in training throughput of ML models. By rapidly programming the server-scale fabric in our hardware testbed, Morphlux can replace a failed accelerator chip with a healthy one in 1.2 seconds. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_03674 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Morphlux: Transforming Torus Fabrics for Efficient Multi-tenant ML Kumar, Abhishek Vijaya Ding, Eric Devraj, Arjun Bunandar, Darius Singh, Rachee Networking and Internet Architecture Hardware Architecture Machine Learning We develop Morphlux, a server-scale programmable photonic fabric to interconnect accelerators within servers. We show that augmenting state-of-the-art torus-based ML data-centers with Morphlux can improve the bandwidth of tenant compute allocations by up to 66%, reduce compute fragmentation by up to 70%, and minimize the blast radius of chip failures. We develop a novel end-to-end hardware prototype of Morphlux to demonstrate these performance benefits which translate to 1.72X improvement in training throughput of ML models. By rapidly programming the server-scale fabric in our hardware testbed, Morphlux can replace a failed accelerator chip with a healthy one in 1.2 seconds. |
| title | Morphlux: Transforming Torus Fabrics for Efficient Multi-tenant ML |
| topic | Networking and Internet Architecture Hardware Architecture Machine Learning |
| url | https://arxiv.org/abs/2508.03674 |