STAGE: A Symbolic Tensor grAph GEnerator for distributed AI system co-design
Fuente:
arXiv
Saved in:
| Main Authors: | Man, Changhai, Park, Joongun, Wu, Hanjiang, Xu, Huan, Sridharan, Srinivas, Krishna, Tushar |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Flint: Compiler Enabled Cluster-Free Design Space Exploration for Distributed ML
by: Yoo, Jinsun, et al.
Published: (2026)
by: Yoo, Jinsun, et al.
Published: (2026)
COSMIC: Enabling Full-Stack Co-Design and Optimization of Distributed Machine Learning Systems
by: Raju, Aditi, et al.
Published: (2025)
by: Raju, Aditi, et al.
Published: (2025)
Characterizing the Efficiency of Distributed Training: A Power, Performance, and Thermal Perspective
by: Go, Seokjin, et al.
Published: (2025)
by: Go, Seokjin, et al.
Published: (2025)
Enhancing Scalability and Performance in Influence Maximization with Optimized Parallel Processing
by: Wu, Hanjiang, et al.
Published: (2024)
by: Wu, Hanjiang, et al.
Published: (2024)
Towards a Standardized Representation for Deep Learning Collective Algorithms
by: Yoo, Jinsun, et al.
Published: (2024)
by: Yoo, Jinsun, et al.
Published: (2024)
LayerDAG: A Layerwise Autoregressive Diffusion Model for Directed Acyclic Graph Generation
by: Li, Mufei, et al.
Published: (2024)
by: Li, Mufei, et al.
Published: (2024)
ASTRA-sim2.0: Modeling Hierarchical Networks and Disaggregated Systems for Large-model Training at Scale
by: Won, William, et al.
Published: (2023)
by: Won, William, et al.
Published: (2023)
MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces
by: Sridharan, Srinivas, et al.
Published: (2026)
by: Sridharan, Srinivas, et al.
Published: (2026)
CELLO: Co-designing Schedule and Hybrid Implicit/Explicit Buffer for Complex Tensor Reuse
by: Garg, Raveesh, et al.
Published: (2023)
by: Garg, Raveesh, et al.
Published: (2023)
Evaluating Cross-Architecture Performance Modeling of Distributed ML Workloads Using StableHLO
by: Svedas, Jonas, et al.
Published: (2026)
by: Svedas, Jonas, et al.
Published: (2026)
LIBRA: Enabling Workload-aware Multi-dimensional Network Topology Optimization for Distributed Training of Large AI Models
by: Won, William, et al.
Published: (2021)
by: Won, William, et al.
Published: (2021)
HARP: A Taxonomy for Heterogeneous and Hierarchical Processors for Mixed-reuse Workloads
by: Garg, Raveesh, et al.
Published: (2025)
by: Garg, Raveesh, et al.
Published: (2025)
KnapsackLB: Enabling Performance-Aware Layer-4 Load Balancing
by: Gandhi, Rohan, et al.
Published: (2024)
by: Gandhi, Rohan, et al.
Published: (2024)
Clock Distribution with Gradient TRIX
by: Lenzen, Christoph, et al.
Published: (2023)
by: Lenzen, Christoph, et al.
Published: (2023)
A distributed system perspective on Backscatter systems: A review
by: Xiao, Tonghuan, et al.
Published: (2025)
by: Xiao, Tonghuan, et al.
Published: (2025)
Scalable Systems and Software Architectures for High-Performance Computing on cloud platforms
by: Ramesh, Risshab Srinivas
Published: (2024)
by: Ramesh, Risshab Srinivas
Published: (2024)
Synergistic Tensor and Pipeline Parallelism
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
Scrutiny new framework in integrated distributed reliable systems
by: Gashti, Mehdi Zekriyapanah
Published: (2025)
by: Gashti, Mehdi Zekriyapanah
Published: (2025)
Federated Single Sign-On and Zero Trust Co-design for AI and HPC Digital Research Infrastructures
by: Alam, Sadaf R., et al.
Published: (2024)
by: Alam, Sadaf R., et al.
Published: (2024)
CAT: Cellular Automata on Tensor cores
by: Navarro, Cristóbal A., et al.
Published: (2024)
by: Navarro, Cristóbal A., et al.
Published: (2024)
TACOS: Topology-Aware Collective Algorithm Synthesizer for Distributed Machine Learning
by: Won, William, et al.
Published: (2023)
by: Won, William, et al.
Published: (2023)
BladeDISC++: Memory Optimizations Based On Symbolic Shape
by: Yuan, Xiulong, et al.
Published: (2024)
by: Yuan, Xiulong, et al.
Published: (2024)
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
by: Zheng, Size, et al.
Published: (2025)
by: Zheng, Size, et al.
Published: (2025)
A monitoring system for collecting and aggregating metrics from distributed clouds
by: Ranković, Tamara, et al.
Published: (2026)
by: Ranković, Tamara, et al.
Published: (2026)
Adaptive Asynchronous Work-Stealing for distributed load-balancing in heterogeneous systems
by: Fernandes, João B., et al.
Published: (2024)
by: Fernandes, João B., et al.
Published: (2024)
Generating representative macrobenchmark microservice systems from distributed traces with Palette
by: Anand, Vaastav, et al.
Published: (2025)
by: Anand, Vaastav, et al.
Published: (2025)
Maya: Optimizing Deep Learning Training Workloads using GPU Runtime Emulation
by: Yarlagadda, Srihas, et al.
Published: (2025)
by: Yarlagadda, Srihas, et al.
Published: (2025)
Accelerating Sparse MTTKRP for Small Tensor Decomposition on GPU
by: Wijeratne, Sasindu, et al.
Published: (2025)
by: Wijeratne, Sasindu, et al.
Published: (2025)
Collaborative Inference Acceleration with Non-Penetrative Tensor Partitioning
by: Liu, Zhibang, et al.
Published: (2025)
by: Liu, Zhibang, et al.
Published: (2025)
Federated Learning Using Coupled Tensor Train Decomposition
by: Zhang, Xiangtao, et al.
Published: (2024)
by: Zhang, Xiangtao, et al.
Published: (2024)
Do We Need Tensor Cores for Stencil Computations?
by: Gu, Qiqi, et al.
Published: (2026)
by: Gu, Qiqi, et al.
Published: (2026)
TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training
by: Liu, Man, et al.
Published: (2026)
by: Liu, Man, et al.
Published: (2026)
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
by: Chen, Haoyu, et al.
Published: (2025)
by: Chen, Haoyu, et al.
Published: (2025)
PRISM: Processing-In-Memory Sparse MTTKRP for Tensor Decomposition Acceleration
by: Pacheco, Daniel, et al.
Published: (2026)
by: Pacheco, Daniel, et al.
Published: (2026)
Accelerating Heterogeneous Tensor Parallelism via Flexible Workload Control
by: Wang, Zhigang, et al.
Published: (2024)
by: Wang, Zhigang, et al.
Published: (2024)
Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism
by: Srivatsa, Vikranth, et al.
Published: (2026)
by: Srivatsa, Vikranth, et al.
Published: (2026)
Accelerating Drug Discovery in AutoDock-GPU with Tensor Cores
by: Schieffer, Gabin, et al.
Published: (2024)
by: Schieffer, Gabin, et al.
Published: (2024)
DynoStore: A wide-area distribution system for the management of data over heterogeneous storage
by: Sanchez-Gallegos, Dante D., et al.
Published: (2025)
by: Sanchez-Gallegos, Dante D., et al.
Published: (2025)
High Performance Unstructured SpMM Computation Using Tensor Cores
by: Okanovic, Patrik, et al.
Published: (2024)
by: Okanovic, Patrik, et al.
Published: (2024)
Predictive Performance of Photonic SRAM-based In-Memory Computing for Tensor Decomposition
by: Wijeratne, Sasindu, et al.
Published: (2025)
by: Wijeratne, Sasindu, et al.
Published: (2025)
Similar Items
-
Flint: Compiler Enabled Cluster-Free Design Space Exploration for Distributed ML
by: Yoo, Jinsun, et al.
Published: (2026) -
COSMIC: Enabling Full-Stack Co-Design and Optimization of Distributed Machine Learning Systems
by: Raju, Aditi, et al.
Published: (2025) -
Characterizing the Efficiency of Distributed Training: A Power, Performance, and Thermal Perspective
by: Go, Seokjin, et al.
Published: (2025) -
Enhancing Scalability and Performance in Influence Maximization with Optimized Parallel Processing
by: Wu, Hanjiang, et al.
Published: (2024) -
Towards a Standardized Representation for Deep Learning Collective Algorithms
by: Yoo, Jinsun, et al.
Published: (2024)