COSMIC: Enabling Full-Stack Co-Design and Optimization of Distributed Machine Learning Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Raju, Aditi, Ni, Jared, Won, William, Man, Changhai, Krishnan, Srivatsan, Sridharan, Srinivas, Yazdanbakhsh, Amir, Krishna, Tushar, Reddi, Vijay Janapa |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Flint: Compiler Enabled Cluster-Free Design Space Exploration for Distributed ML
von: Yoo, Jinsun, et al.
Veröffentlicht: (2026)
von: Yoo, Jinsun, et al.
Veröffentlicht: (2026)
STAGE: A Symbolic Tensor grAph GEnerator for distributed AI system co-design
von: Man, Changhai, et al.
Veröffentlicht: (2025)
von: Man, Changhai, et al.
Veröffentlicht: (2025)
Towards a Standardized Representation for Deep Learning Collective Algorithms
von: Yoo, Jinsun, et al.
Veröffentlicht: (2024)
von: Yoo, Jinsun, et al.
Veröffentlicht: (2024)
ASTRA-sim2.0: Modeling Hierarchical Networks and Disaggregated Systems for Large-model Training at Scale
von: Won, William, et al.
Veröffentlicht: (2023)
von: Won, William, et al.
Veröffentlicht: (2023)
FedStaleWeight: Buffered Asynchronous Federated Learning with Fair Aggregation via Staleness Reweighting
von: Ma, Jeffrey, et al.
Veröffentlicht: (2024)
von: Ma, Jeffrey, et al.
Veröffentlicht: (2024)
LayerDAG: A Layerwise Autoregressive Diffusion Model for Directed Acyclic Graph Generation
von: Li, Mufei, et al.
Veröffentlicht: (2024)
von: Li, Mufei, et al.
Veröffentlicht: (2024)
MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces
von: Sridharan, Srinivas, et al.
Veröffentlicht: (2026)
von: Sridharan, Srinivas, et al.
Veröffentlicht: (2026)
Evaluating Cross-Architecture Performance Modeling of Distributed ML Workloads Using StableHLO
von: Svedas, Jonas, et al.
Veröffentlicht: (2026)
von: Svedas, Jonas, et al.
Veröffentlicht: (2026)
LIBRA: Enabling Workload-aware Multi-dimensional Network Topology Optimization for Distributed Training of Large AI Models
von: Won, William, et al.
Veröffentlicht: (2021)
von: Won, William, et al.
Veröffentlicht: (2021)
SwizzlePerf: Hardware-Aware LLMs for GPU Kernel Performance Optimization
von: Tschand, Arya, et al.
Veröffentlicht: (2025)
von: Tschand, Arya, et al.
Veröffentlicht: (2025)
CELLO: Co-designing Schedule and Hybrid Implicit/Explicit Buffer for Complex Tensor Reuse
von: Garg, Raveesh, et al.
Veröffentlicht: (2023)
von: Garg, Raveesh, et al.
Veröffentlicht: (2023)
KnapsackLB: Enabling Performance-Aware Layer-4 Load Balancing
von: Gandhi, Rohan, et al.
Veröffentlicht: (2024)
von: Gandhi, Rohan, et al.
Veröffentlicht: (2024)
TACOS: Topology-Aware Collective Algorithm Synthesizer for Distributed Machine Learning
von: Won, William, et al.
Veröffentlicht: (2023)
von: Won, William, et al.
Veröffentlicht: (2023)
Persistent HyTM via Fast Path Fine-Grained Locking
von: Coccimiglio, Gaetano, et al.
Veröffentlicht: (2025)
von: Coccimiglio, Gaetano, et al.
Veröffentlicht: (2025)
How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving
von: Wu, Hanjiang, et al.
Veröffentlicht: (2026)
von: Wu, Hanjiang, et al.
Veröffentlicht: (2026)
MIST: A Co-Design Framework for Heterogeneous, Multi-Stage LLM Inference
von: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Veröffentlicht: (2025)
von: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Veröffentlicht: (2025)
HARP: A Taxonomy for Heterogeneous and Hierarchical Processors for Mixed-reuse Workloads
von: Garg, Raveesh, et al.
Veröffentlicht: (2025)
von: Garg, Raveesh, et al.
Veröffentlicht: (2025)
Enabling an OpenStack-based cloud on top of RISC-V hardware
von: Marrón, Diego, et al.
Veröffentlicht: (2024)
von: Marrón, Diego, et al.
Veröffentlicht: (2024)
Training Foundation Models on a Full-Stack AMD Platform: Compute, Networking, and System Design
von: Anthony, Quentin, et al.
Veröffentlicht: (2025)
von: Anthony, Quentin, et al.
Veröffentlicht: (2025)
Blockchain Transaction Conflicts: A Historical Perspective
von: Anjana, Parwat Singh, et al.
Veröffentlicht: (2025)
von: Anjana, Parwat Singh, et al.
Veröffentlicht: (2025)
Stingray: Fast Concurrent Transactions Without Consensus
von: Sridhar, Srivatsan, et al.
Veröffentlicht: (2025)
von: Sridhar, Srivatsan, et al.
Veröffentlicht: (2025)
Clock Distribution with Gradient TRIX
von: Lenzen, Christoph, et al.
Veröffentlicht: (2023)
von: Lenzen, Christoph, et al.
Veröffentlicht: (2023)
Scaling on Frontier: Uncertainty Quantification Workflow Applications using ExaWorks to Enable Full System Utilization
von: Titov, Mikhail, et al.
Veröffentlicht: (2024)
von: Titov, Mikhail, et al.
Veröffentlicht: (2024)
Scalable Systems and Software Architectures for High-Performance Computing on cloud platforms
von: Ramesh, Risshab Srinivas
Veröffentlicht: (2024)
von: Ramesh, Risshab Srinivas
Veröffentlicht: (2024)
Efficient Parallel Execution of Blockchain Transactions Leveraging Conflict Specifications
von: Anjana, Parwat Singh, et al.
Veröffentlicht: (2025)
von: Anjana, Parwat Singh, et al.
Veröffentlicht: (2025)
PISA: An Adversarial Approach To Comparing Task Graph Scheduling Algorithms
von: Coleman, Jared, et al.
Veröffentlicht: (2024)
von: Coleman, Jared, et al.
Veröffentlicht: (2024)
On Software Ageing Indicators in OpenStack
von: Yazvinskyi, Yevhen, et al.
Veröffentlicht: (2024)
von: Yazvinskyi, Yevhen, et al.
Veröffentlicht: (2024)
Maya: Optimizing Deep Learning Training Workloads using GPU Runtime Emulation
von: Yarlagadda, Srihas, et al.
Veröffentlicht: (2025)
von: Yarlagadda, Srihas, et al.
Veröffentlicht: (2025)
DRust: Language-Guided Distributed Shared Memory with Fine Granularity, Full Transparency, and Ultra Efficiency
von: Ma, Haoran, et al.
Veröffentlicht: (2024)
von: Ma, Haoran, et al.
Veröffentlicht: (2024)
Nakamoto Consensus under Bounded Processing Capacity
von: Kiffer, Lucianna, et al.
Veröffentlicht: (2023)
von: Kiffer, Lucianna, et al.
Veröffentlicht: (2023)
Consensus Under Adversary Majority Done Right
von: Sridhar, Srivatsan, et al.
Veröffentlicht: (2024)
von: Sridhar, Srivatsan, et al.
Veröffentlicht: (2024)
A common parallel framework for LLP combinatorial problems
von: Alves, David Ribeiro, et al.
Veröffentlicht: (2026)
von: Alves, David Ribeiro, et al.
Veröffentlicht: (2026)
SWARM+: Scalable and Resilient Multi-Agent Consensus for Fully-Decentralized Data-Aware Workload Management
von: Thareja, Komal, et al.
Veröffentlicht: (2026)
von: Thareja, Komal, et al.
Veröffentlicht: (2026)
A Full Stack Framework for High Performance Quantum-Classical Computing
von: Zhan, Xin, et al.
Veröffentlicht: (2025)
von: Zhan, Xin, et al.
Veröffentlicht: (2025)
Parameterized Task Graph Scheduling Algorithm for Comparing Algorithmic Components
von: Coleman, Jared, et al.
Veröffentlicht: (2024)
von: Coleman, Jared, et al.
Veröffentlicht: (2024)
KUBEDIRECT: Unleashing the Full Power of the Cluster Manager for Serverless Computing
von: Qi, Sheng, et al.
Veröffentlicht: (2026)
von: Qi, Sheng, et al.
Veröffentlicht: (2026)
Characterizing the Efficiency of Distributed Training: A Power, Performance, and Thermal Perspective
von: Go, Seokjin, et al.
Veröffentlicht: (2025)
von: Go, Seokjin, et al.
Veröffentlicht: (2025)
On the Convergence of Malleability and the HPC PowerStack: Exploiting Dynamism in Over-Provisioned and Power-Constrained HPC Systems
von: Arima, Eishi, et al.
Veröffentlicht: (2024)
von: Arima, Eishi, et al.
Veröffentlicht: (2024)
Fix: externalizing network I/O in serverless computing
von: Deng, Yuhan, et al.
Veröffentlicht: (2025)
von: Deng, Yuhan, et al.
Veröffentlicht: (2025)
Beyond Pre-Training: The Full Lifecycle of Foundation Models on HPC Systems
von: Conciatore, Dino, et al.
Veröffentlicht: (2026)
von: Conciatore, Dino, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Flint: Compiler Enabled Cluster-Free Design Space Exploration for Distributed ML
von: Yoo, Jinsun, et al.
Veröffentlicht: (2026) -
STAGE: A Symbolic Tensor grAph GEnerator for distributed AI system co-design
von: Man, Changhai, et al.
Veröffentlicht: (2025) -
Towards a Standardized Representation for Deep Learning Collective Algorithms
von: Yoo, Jinsun, et al.
Veröffentlicht: (2024) -
ASTRA-sim2.0: Modeling Hierarchical Networks and Disaggregated Systems for Large-model Training at Scale
von: Won, William, et al.
Veröffentlicht: (2023) -
FedStaleWeight: Buffered Asynchronous Federated Learning with Fair Aggregation via Staleness Reweighting
von: Ma, Jeffrey, et al.
Veröffentlicht: (2024)