Efficient Direct-Connect Topologies for Collective Communications
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2022
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866915131707883520 |
|---|---|
| author | Zhao, Liangyu Pal, Siddharth Chugh, Tapan Wang, Weiyang Fantl, Jason Basu, Prithwish Khoury, Joud Krishnamurthy, Arvind |
| author_facet | Zhao, Liangyu Pal, Siddharth Chugh, Tapan Wang, Weiyang Fantl, Jason Basu, Prithwish Khoury, Joud Krishnamurthy, Arvind |
| contents | We consider the problem of distilling efficient network topologies for collective communications. We provide an algorithmic framework for constructing direct-connect topologies optimized for the latency vs. bandwidth trade-off associated with the workload. Our approach synthesizes many different topologies and schedules for a given cluster size and degree and then identifies the appropriate topology and schedule for a given workload. Our algorithms start from small, optimal base topologies and associated communication schedules and use techniques that can be iteratively applied to derive much larger topologies and schedules. Additionally, we incorporate well-studied large-scale graph topologies into our algorithmic framework by producing efficient collective schedules for them using a novel polynomial-time algorithm. Our evaluation uses multiple testbeds and large-scale simulations to demonstrate significant performance benefits from our derived topologies and schedules. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2202_03356 |
| institution | arXiv |
| publishDate | 2022 |
| record_format | arxiv |
| spellingShingle | Efficient Direct-Connect Topologies for Collective Communications Zhao, Liangyu Pal, Siddharth Chugh, Tapan Wang, Weiyang Fantl, Jason Basu, Prithwish Khoury, Joud Krishnamurthy, Arvind Networking and Internet Architecture Distributed, Parallel, and Cluster Computing Machine Learning We consider the problem of distilling efficient network topologies for collective communications. We provide an algorithmic framework for constructing direct-connect topologies optimized for the latency vs. bandwidth trade-off associated with the workload. Our approach synthesizes many different topologies and schedules for a given cluster size and degree and then identifies the appropriate topology and schedule for a given workload. Our algorithms start from small, optimal base topologies and associated communication schedules and use techniques that can be iteratively applied to derive much larger topologies and schedules. Additionally, we incorporate well-studied large-scale graph topologies into our algorithmic framework by producing efficient collective schedules for them using a novel polynomial-time algorithm. Our evaluation uses multiple testbeds and large-scale simulations to demonstrate significant performance benefits from our derived topologies and schedules. |
| title | Efficient Direct-Connect Topologies for Collective Communications |
| topic | Networking and Internet Architecture Distributed, Parallel, and Cluster Computing Machine Learning |
| url | https://arxiv.org/abs/2202.03356 |