Flint: Compiler Enabled Cluster-Free Design Space Exploration for Distributed ML
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yoo, Jinsun, Cowan, Meghan, Du, Zheng, Man, Changhai, Sridharan, Srinivas, Krishna, Tushar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards a Standardized Representation for Deep Learning Collective Algorithms
von: Yoo, Jinsun, et al.
Veröffentlicht: (2024)
von: Yoo, Jinsun, et al.
Veröffentlicht: (2024)
COSMIC: Enabling Full-Stack Co-Design and Optimization of Distributed Machine Learning Systems
von: Raju, Aditi, et al.
Veröffentlicht: (2025)
von: Raju, Aditi, et al.
Veröffentlicht: (2025)
STAGE: A Symbolic Tensor grAph GEnerator for distributed AI system co-design
von: Man, Changhai, et al.
Veröffentlicht: (2025)
von: Man, Changhai, et al.
Veröffentlicht: (2025)
Evaluating Cross-Architecture Performance Modeling of Distributed ML Workloads Using StableHLO
von: Svedas, Jonas, et al.
Veröffentlicht: (2026)
von: Svedas, Jonas, et al.
Veröffentlicht: (2026)
LayerDAG: A Layerwise Autoregressive Diffusion Model for Directed Acyclic Graph Generation
von: Li, Mufei, et al.
Veröffentlicht: (2024)
von: Li, Mufei, et al.
Veröffentlicht: (2024)
ASTRA-sim2.0: Modeling Hierarchical Networks and Disaggregated Systems for Large-model Training at Scale
von: Won, William, et al.
Veröffentlicht: (2023)
von: Won, William, et al.
Veröffentlicht: (2023)
DeepCompile: A Compiler-Driven Approach to Optimizing Distributed Deep Learning Training
von: Tanaka, Masahiro, et al.
Veröffentlicht: (2025)
von: Tanaka, Masahiro, et al.
Veröffentlicht: (2025)
LIBRA: Enabling Workload-aware Multi-dimensional Network Topology Optimization for Distributed Training of Large AI Models
von: Won, William, et al.
Veröffentlicht: (2021)
von: Won, William, et al.
Veröffentlicht: (2021)
Clock Distribution with Gradient TRIX
von: Lenzen, Christoph, et al.
Veröffentlicht: (2023)
von: Lenzen, Christoph, et al.
Veröffentlicht: (2023)
KnapsackLB: Enabling Performance-Aware Layer-4 Load Balancing
von: Gandhi, Rohan, et al.
Veröffentlicht: (2024)
von: Gandhi, Rohan, et al.
Veröffentlicht: (2024)
COMET: A Comprehensive Cluster Design Methodology for Distributed Deep Learning Training
von: Kadiyala, Divya Kiran, et al.
Veröffentlicht: (2022)
von: Kadiyala, Divya Kiran, et al.
Veröffentlicht: (2022)
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
von: Zheng, Size, et al.
Veröffentlicht: (2025)
von: Zheng, Size, et al.
Veröffentlicht: (2025)
Optimizing Compilation for Distributed Quantum Computing via Clustering and Annealing
von: Zhou, Ruilin, et al.
Veröffentlicht: (2025)
von: Zhou, Ruilin, et al.
Veröffentlicht: (2025)
Designing Dense Satellite Clusters for Distributed Space-based Datacenters
von: Pénot, Jules, et al.
Veröffentlicht: (2026)
von: Pénot, Jules, et al.
Veröffentlicht: (2026)
MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces
von: Sridharan, Srinivas, et al.
Veröffentlicht: (2026)
von: Sridharan, Srinivas, et al.
Veröffentlicht: (2026)
PAL: A Variability-Aware Policy for Scheduling ML Workloads in GPU Clusters
von: Jain, Rutwik, et al.
Veröffentlicht: (2024)
von: Jain, Rutwik, et al.
Veröffentlicht: (2024)
TACOS: Topology-Aware Collective Algorithm Synthesizer for Distributed Machine Learning
von: Won, William, et al.
Veröffentlicht: (2023)
von: Won, William, et al.
Veröffentlicht: (2023)
Towards Easy and Realistic Network Infrastructure Testing for Large-scale Machine Learning
von: Yoo, Jinsun, et al.
Veröffentlicht: (2025)
von: Yoo, Jinsun, et al.
Veröffentlicht: (2025)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
von: Zhang, WenZheng, et al.
Veröffentlicht: (2024)
von: Zhang, WenZheng, et al.
Veröffentlicht: (2024)
Deterministic Collision-Free Exploration of Unknown Anonymous Graphs
von: Bhagat, Subhash, et al.
Veröffentlicht: (2024)
von: Bhagat, Subhash, et al.
Veröffentlicht: (2024)
HARP: A Taxonomy for Heterogeneous and Hierarchical Processors for Mixed-reuse Workloads
von: Garg, Raveesh, et al.
Veröffentlicht: (2025)
von: Garg, Raveesh, et al.
Veröffentlicht: (2025)
PowerTrip: Exploiting Federated Heterogeneous Datacenter Power for Distributed ML Training
von: Mehboob, Talha, et al.
Veröffentlicht: (2025)
von: Mehboob, Talha, et al.
Veröffentlicht: (2025)
NetSenseML: Network-Adaptive Compression for Efficient Distributed Machine Learning
von: Wang, Yisu, et al.
Veröffentlicht: (2025)
von: Wang, Yisu, et al.
Veröffentlicht: (2025)
DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators
von: Mo, Zhiwen, et al.
Veröffentlicht: (2026)
von: Mo, Zhiwen, et al.
Veröffentlicht: (2026)
HPCTransCompile: An AI Compiler Generated Dataset for High-Performance CUDA Transpilation and LLM Preliminary Exploration
von: Lv, Jiaqi, et al.
Veröffentlicht: (2025)
von: Lv, Jiaqi, et al.
Veröffentlicht: (2025)
A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via Faro
von: Jeon, Beomyeol, et al.
Veröffentlicht: (2024)
von: Jeon, Beomyeol, et al.
Veröffentlicht: (2024)
DiT-HC: Enabling Efficient Training of Visual Generation Model DiT on HPC-oriented CPU Cluster
von: Zhang, Jinxiao, et al.
Veröffentlicht: (2026)
von: Zhang, Jinxiao, et al.
Veröffentlicht: (2026)
TokenSim: Enabling Hardware and Software Exploration for Large Language Model Inference Systems
von: Wu, Feiyang, et al.
Veröffentlicht: (2025)
von: Wu, Feiyang, et al.
Veröffentlicht: (2025)
Optimizing Memory Allocation in Distributed Clusters with Predictive Modeling
von: Bader, Jonathan, et al.
Veröffentlicht: (2026)
von: Bader, Jonathan, et al.
Veröffentlicht: (2026)
Solutions for Distributed Memory Access Mechanism on HPC Clusters
von: Meizner, Jan, et al.
Veröffentlicht: (2025)
von: Meizner, Jan, et al.
Veröffentlicht: (2025)
CarbonFlex: Enabling Carbon-aware Provisioning and Scheduling for Cloud Clusters
von: Hanafy, Walid A., et al.
Veröffentlicht: (2025)
von: Hanafy, Walid A., et al.
Veröffentlicht: (2025)
An Explorative Study on Distributed Computing Techniques in Training and Inference of Large Language Models
von: Hakim, Sheikh Azizul, et al.
Veröffentlicht: (2025)
von: Hakim, Sheikh Azizul, et al.
Veröffentlicht: (2025)
HPAC-ML: A Programming Model for Embedding ML Surrogates in Scientific Applications
von: Fink, Zane, et al.
Veröffentlicht: (2024)
von: Fink, Zane, et al.
Veröffentlicht: (2024)
CELLO: Co-designing Schedule and Hybrid Implicit/Explicit Buffer for Complex Tensor Reuse
von: Garg, Raveesh, et al.
Veröffentlicht: (2023)
von: Garg, Raveesh, et al.
Veröffentlicht: (2023)
MLDSE: Scaling Design Space Exploration Infrastructure for Multi-Level Hardware
von: Qu, Huanyu, et al.
Veröffentlicht: (2025)
von: Qu, Huanyu, et al.
Veröffentlicht: (2025)
Parallel Online Directed Acyclic Graph Exploration for Atlasing Soft-Matter Assembly Configuration Spaces
von: Prabhu, Rahul, et al.
Veröffentlicht: (2024)
von: Prabhu, Rahul, et al.
Veröffentlicht: (2024)
How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving
von: Wu, Hanjiang, et al.
Veröffentlicht: (2026)
von: Wu, Hanjiang, et al.
Veröffentlicht: (2026)
Characterizing the Efficiency of Distributed Training: A Power, Performance, and Thermal Perspective
von: Go, Seokjin, et al.
Veröffentlicht: (2025)
von: Go, Seokjin, et al.
Veröffentlicht: (2025)
QONNECT: A QoS-Aware Orchestration System for Distributed Kubernetes Clusters
von: Aslan, Haci Ismail, et al.
Veröffentlicht: (2025)
von: Aslan, Haci Ismail, et al.
Veröffentlicht: (2025)
EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
von: Du, Jiangsu, et al.
Veröffentlicht: (2025)
von: Du, Jiangsu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Towards a Standardized Representation for Deep Learning Collective Algorithms
von: Yoo, Jinsun, et al.
Veröffentlicht: (2024) -
COSMIC: Enabling Full-Stack Co-Design and Optimization of Distributed Machine Learning Systems
von: Raju, Aditi, et al.
Veröffentlicht: (2025) -
STAGE: A Symbolic Tensor grAph GEnerator for distributed AI system co-design
von: Man, Changhai, et al.
Veröffentlicht: (2025) -
Evaluating Cross-Architecture Performance Modeling of Distributed ML Workloads Using StableHLO
von: Svedas, Jonas, et al.
Veröffentlicht: (2026) -
LayerDAG: A Layerwise Autoregressive Diffusion Model for Directed Acyclic Graph Generation
von: Li, Mufei, et al.
Veröffentlicht: (2024)