Towards a Standardized Representation for Deep Learning Collective Algorithms
Fuente:
arXiv
Saved in:
| Main Authors: | Yoo, Jinsun, Won, William, Cowan, Meghan, Jiang, Nan, Klenk, Benjamin, Sridharan, Srinivas, Krishna, Tushar |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Flint: Compiler Enabled Cluster-Free Design Space Exploration for Distributed ML
by: Yoo, Jinsun, et al.
Published: (2026)
by: Yoo, Jinsun, et al.
Published: (2026)
ASTRA-sim2.0: Modeling Hierarchical Networks and Disaggregated Systems for Large-model Training at Scale
by: Won, William, et al.
Published: (2023)
by: Won, William, et al.
Published: (2023)
COSMIC: Enabling Full-Stack Co-Design and Optimization of Distributed Machine Learning Systems
by: Raju, Aditi, et al.
Published: (2025)
by: Raju, Aditi, et al.
Published: (2025)
TACOS: Topology-Aware Collective Algorithm Synthesizer for Distributed Machine Learning
by: Won, William, et al.
Published: (2023)
by: Won, William, et al.
Published: (2023)
STAGE: A Symbolic Tensor grAph GEnerator for distributed AI system co-design
by: Man, Changhai, et al.
Published: (2025)
by: Man, Changhai, et al.
Published: (2025)
LIBRA: Enabling Workload-aware Multi-dimensional Network Topology Optimization for Distributed Training of Large AI Models
by: Won, William, et al.
Published: (2021)
by: Won, William, et al.
Published: (2021)
Towards Easy and Realistic Network Infrastructure Testing for Large-scale Machine Learning
by: Yoo, Jinsun, et al.
Published: (2025)
by: Yoo, Jinsun, et al.
Published: (2025)
MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces
by: Sridharan, Srinivas, et al.
Published: (2026)
by: Sridharan, Srinivas, et al.
Published: (2026)
Maya: Optimizing Deep Learning Training Workloads using GPU Runtime Emulation
by: Yarlagadda, Srihas, et al.
Published: (2025)
by: Yarlagadda, Srihas, et al.
Published: (2025)
LayerDAG: A Layerwise Autoregressive Diffusion Model for Directed Acyclic Graph Generation
by: Li, Mufei, et al.
Published: (2024)
by: Li, Mufei, et al.
Published: (2024)
A Tabular Schedule Abstraction for Communication-Aware Evaluation of Pipeline-Parallel LLM Training
by: Barley, Daniel, et al.
Published: (2026)
by: Barley, Daniel, et al.
Published: (2026)
Towards the Democratization and Standardization of Dynamic Resources with MPI Spawning
by: Iserte, Sergio, et al.
Published: (2026)
by: Iserte, Sergio, et al.
Published: (2026)
HARP: A Taxonomy for Heterogeneous and Hierarchical Processors for Mixed-reuse Workloads
by: Garg, Raveesh, et al.
Published: (2025)
by: Garg, Raveesh, et al.
Published: (2025)
KnapsackLB: Enabling Performance-Aware Layer-4 Load Balancing
by: Gandhi, Rohan, et al.
Published: (2024)
by: Gandhi, Rohan, et al.
Published: (2024)
Clock Distribution with Gradient TRIX
by: Lenzen, Christoph, et al.
Published: (2023)
by: Lenzen, Christoph, et al.
Published: (2023)
Towards Optimal Deterministic LOCAL Algorithms on Trees
by: Brandt, Sebastian, et al.
Published: (2025)
by: Brandt, Sebastian, et al.
Published: (2025)
Evaluating Cross-Architecture Performance Modeling of Distributed ML Workloads Using StableHLO
by: Svedas, Jonas, et al.
Published: (2026)
by: Svedas, Jonas, et al.
Published: (2026)
Selection of Supervised Learning-based Sparse Matrix Reordering Algorithms
by: Tang, Tao, et al.
Published: (2025)
by: Tang, Tao, et al.
Published: (2025)
CELLO: Co-designing Schedule and Hybrid Implicit/Explicit Buffer for Complex Tensor Reuse
by: Garg, Raveesh, et al.
Published: (2023)
by: Garg, Raveesh, et al.
Published: (2023)
RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
by: Brock, Benjamin, et al.
Published: (2023)
by: Brock, Benjamin, et al.
Published: (2023)
HPC Containers for EBRAINS: Towards Portable Cross-Domain Software Environment
by: Singh, Krishna Kant, et al.
Published: (2026)
by: Singh, Krishna Kant, et al.
Published: (2026)
AgentServe: Algorithm-System Co-Design for Efficient Agentic AI Serving on a Consumer-Grade GPU
by: Zhang, Yuning, et al.
Published: (2026)
by: Zhang, Yuning, et al.
Published: (2026)
Scalable Systems and Software Architectures for High-Performance Computing on cloud platforms
by: Ramesh, Risshab Srinivas
Published: (2024)
by: Ramesh, Risshab Srinivas
Published: (2024)
Deep Learning-Enabled Supercritical Flame Simulation at Detailed Chemistry and Real-Fluid Accuracy Towards Trillion-Cell Scale
by: Guo, Zhuoqiang, et al.
Published: (2025)
by: Guo, Zhuoqiang, et al.
Published: (2025)
AgentX: Towards Orchestrating Robust Agentic Workflow Patterns with FaaS-hosted MCP Services
by: Tokal, Shiva Sai Krishna Anand, et al.
Published: (2025)
by: Tokal, Shiva Sai Krishna Anand, et al.
Published: (2025)
HiCCL: A Hierarchical Collective Communication Library
by: Hidayetoglu, Mert, et al.
Published: (2024)
by: Hidayetoglu, Mert, et al.
Published: (2024)
MicroPython Testbed for Federated Learning Algorithms
by: Popovic, Miroslav, et al.
Published: (2024)
by: Popovic, Miroslav, et al.
Published: (2024)
A Seesaw Model Attack Algorithm for Distributed Learning
by: Yang, Kun, et al.
Published: (2024)
by: Yang, Kun, et al.
Published: (2024)
Developing Elementary Federated Learning Algorithms Leveraging the ChatGPT
by: Popovic, Miroslav, et al.
Published: (2023)
by: Popovic, Miroslav, et al.
Published: (2023)
DeepCompile: A Compiler-Driven Approach to Optimizing Distributed Deep Learning Training
by: Tanaka, Masahiro, et al.
Published: (2025)
by: Tanaka, Masahiro, et al.
Published: (2025)
COMET: A Comprehensive Cluster Design Methodology for Distributed Deep Learning Training
by: Kadiyala, Divya Kiran, et al.
Published: (2022)
by: Kadiyala, Divya Kiran, et al.
Published: (2022)
On-the-fly Communication-and-Computing to Enable Representation Learning for Distributed Point Clouds
by: Chen, Xu, et al.
Published: (2024)
by: Chen, Xu, et al.
Published: (2024)
DeepVM: Integrating Spot and On-Demand VMs for Cost-Efficient Deep Learning Clusters in the Cloud
by: Kim, Yoochan, et al.
Published: (2024)
by: Kim, Yoochan, et al.
Published: (2024)
EPIC: Abstraction and Polymorphism of In-Network Collectives on Ethernet
by: Yuan, Yitao, et al.
Published: (2026)
by: Yuan, Yitao, et al.
Published: (2026)
Distributed Ranges: A Model for Distributed Data Structures, Algorithms, and Views
by: Brock, Benjamin, et al.
Published: (2024)
by: Brock, Benjamin, et al.
Published: (2024)
AAPA: An Archetype-Aware Predictive Autoscaler with Uncertainty Quantification for Serverless Workloads on Kubernetes
by: Zhang, Guilin, et al.
Published: (2025)
by: Zhang, Guilin, et al.
Published: (2025)
Joint Temporal-Structural Representation Learning for Distributed Fault Discrimination in Microservice Architectures
by: Xue, Yihan, et al.
Published: (2026)
by: Xue, Yihan, et al.
Published: (2026)
HPX -- An open source C++ Standard Library for Parallelism and Concurrency
by: Heller, Thomas, et al.
Published: (2023)
by: Heller, Thomas, et al.
Published: (2023)
Slicing Is All You Need: Towards A Universal One-Sided Algorithm for Distributed Matrix Multiplication
by: Brock, Benjamin, et al.
Published: (2025)
by: Brock, Benjamin, et al.
Published: (2025)
Hiding Latencies in Network-Based Image Loading for Deep Learning
by: Versaci, Francesco, et al.
Published: (2025)
by: Versaci, Francesco, et al.
Published: (2025)
Similar Items
-
Flint: Compiler Enabled Cluster-Free Design Space Exploration for Distributed ML
by: Yoo, Jinsun, et al.
Published: (2026) -
ASTRA-sim2.0: Modeling Hierarchical Networks and Disaggregated Systems for Large-model Training at Scale
by: Won, William, et al.
Published: (2023) -
COSMIC: Enabling Full-Stack Co-Design and Optimization of Distributed Machine Learning Systems
by: Raju, Aditi, et al.
Published: (2025) -
TACOS: Topology-Aware Collective Algorithm Synthesizer for Distributed Machine Learning
by: Won, William, et al.
Published: (2023) -
STAGE: A Symbolic Tensor grAph GEnerator for distributed AI system co-design
by: Man, Changhai, et al.
Published: (2025)