CubicML: Automated ML for Large ML Systems Co-design with ML Prediction of Performance
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wen, Wei, Zhu, Quanyu, Chu, Weiwei, Chen, Wen-Yen, Yang, Jiyan |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Compass: A Decentralized Scheduler for Latency-Sensitive ML Workflows
par: Yang, Yuting, et autres
Publié: (2024)
par: Yang, Yuting, et autres
Publié: (2024)
Cloudless-Training: A Framework to Improve Efficiency of Geo-Distributed ML Training
par: Tan, Wenting, et autres
Publié: (2023)
par: Tan, Wenting, et autres
Publié: (2023)
Scalable Runtime Architecture for Data-driven, Hybrid HPC and ML Workflow Applications
par: Merzky, Andre, et autres
Publié: (2025)
par: Merzky, Andre, et autres
Publié: (2025)
HPAC-ML: A Programming Model for Embedding ML Surrogates in Scientific Applications
par: Fink, Zane, et autres
Publié: (2024)
par: Fink, Zane, et autres
Publié: (2024)
On-device Online Learning and Semantic Management of TinyML Systems
par: Ren, Haoyu, et autres
Publié: (2024)
par: Ren, Haoyu, et autres
Publié: (2024)
Optimizing Split Learning Latency in TinyML-Based IoT Systems
par: Jenhani, Zied, et autres
Publié: (2025)
par: Jenhani, Zied, et autres
Publié: (2025)
TrainMover: An Interruption-Resilient Runtime for ML Training
par: Lao, ChonLam, et autres
Publié: (2024)
par: Lao, ChonLam, et autres
Publié: (2024)
Declarative Data Pipeline for Large Scale ML Services
par: Yang, Yunzhao, et autres
Publié: (2025)
par: Yang, Yunzhao, et autres
Publié: (2025)
ML-based Modeling to Predict I/O Performance on Different Storage Sub-systems
par: Xu, Yiheng, et autres
Publié: (2023)
par: Xu, Yiheng, et autres
Publié: (2023)
GetBatch: Distributed Multi-Object Retrieval for ML Data Loading
par: Aizman, Alex, et autres
Publié: (2026)
par: Aizman, Alex, et autres
Publié: (2026)
ML-based Adaptive Prefetching and Data Placement for US HEP Systems
par: Karanam, Venkat Sai Suman Lamba, et autres
Publié: (2025)
par: Karanam, Venkat Sai Suman Lamba, et autres
Publié: (2025)
A Performance Analyzer for a Public Cloud's ML-Augmented VM Allocator
par: Bostandoost, Roozbeh, et autres
Publié: (2025)
par: Bostandoost, Roozbeh, et autres
Publié: (2025)
ML-ECS: A Collaborative Multimodal Learning Framework for Edge-Cloud Synergies
par: Liu, Yuze, et autres
Publié: (2026)
par: Liu, Yuze, et autres
Publié: (2026)
Loki: A System for Serving ML Inference Pipelines with Hardware and Accuracy Scaling
par: Ahmad, Sohaib, et autres
Publié: (2024)
par: Ahmad, Sohaib, et autres
Publié: (2024)
Evaluating Cross-Architecture Performance Modeling of Distributed ML Workloads Using StableHLO
par: Svedas, Jonas, et autres
Publié: (2026)
par: Svedas, Jonas, et autres
Publié: (2026)
PCCL: Photonic circuit-switched collective communication for distributed ML
par: Kumar, Abhishek Vijaya, et autres
Publié: (2025)
par: Kumar, Abhishek Vijaya, et autres
Publié: (2025)
NetSenseML: Network-Adaptive Compression for Efficient Distributed Machine Learning
par: Wang, Yisu, et autres
Publié: (2025)
par: Wang, Yisu, et autres
Publié: (2025)
PAL: A Variability-Aware Policy for Scheduling ML Workloads in GPU Clusters
par: Jain, Rutwik, et autres
Publié: (2024)
par: Jain, Rutwik, et autres
Publié: (2024)
PowerTrip: Exploiting Federated Heterogeneous Datacenter Power for Distributed ML Training
par: Mehboob, Talha, et autres
Publié: (2025)
par: Mehboob, Talha, et autres
Publié: (2025)
Flint: Compiler Enabled Cluster-Free Design Space Exploration for Distributed ML
par: Yoo, Jinsun, et autres
Publié: (2026)
par: Yoo, Jinsun, et autres
Publié: (2026)
Reimagining RDMA Through the Lens of ML
par: Warraich, Ertza, et autres
Publié: (2025)
par: Warraich, Ertza, et autres
Publié: (2025)
Snowpark: Performant, Secure, User-Friendly Data Engineering and AI/ML Next To Your Data
par: Baker, Brandon, et autres
Publié: (2025)
par: Baker, Brandon, et autres
Publié: (2025)
An Efficient Heterogeneous Co-Design for Fine-Tuning on a Single GPU
par: Yang, Ruijia, et autres
Publié: (2026)
par: Yang, Ruijia, et autres
Publié: (2026)
OptiML: An End-to-End Framework for Program Synthesis and CUDA Kernel Optimization
par: Bhattacharjee, Arijit, et autres
Publié: (2026)
par: Bhattacharjee, Arijit, et autres
Publié: (2026)
Deploying Graph Neural Networks in Wireless Networks: A Link Stability Viewpoint
par: Li, Jun, et autres
Publié: (2024)
par: Li, Jun, et autres
Publié: (2024)
yProv4ML: Effortless Provenance Tracking for Machine Learning Systems
par: Padovani, Gabriele, et autres
Publié: (2025)
par: Padovani, Gabriele, et autres
Publié: (2025)
Magneton: Optimizing Energy Efficiency of ML Systems via Differential Energy Debugging
par: Pan, Yi, et autres
Publié: (2025)
par: Pan, Yi, et autres
Publié: (2025)
tf.data service: A Case for Disaggregating ML Input Data Processing
par: Audibert, Andrew, et autres
Publié: (2022)
par: Audibert, Andrew, et autres
Publié: (2022)
Agentic Operator Generation for ML ASICs
par: Hammond, Alec M., et autres
Publié: (2025)
par: Hammond, Alec M., et autres
Publié: (2025)
A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via Faro
par: Jeon, Beomyeol, et autres
Publié: (2024)
par: Jeon, Beomyeol, et autres
Publié: (2024)
Optimizing Distributed ML Communication with Fused Computation-Collective Operations
par: Punniyamurthy, Kishore, et autres
Publié: (2023)
par: Punniyamurthy, Kishore, et autres
Publié: (2023)
Taming the Memory Beast: Strategies for Reliable ML Training on Kubernetes
par: Ray, Jaideep
Publié: (2024)
par: Ray, Jaideep
Publié: (2024)
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
par: Agrawal, Anirudha, et autres
Publié: (2024)
par: Agrawal, Anirudha, et autres
Publié: (2024)
A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators
par: Colagrande, Luca, et autres
Publié: (2026)
par: Colagrande, Luca, et autres
Publié: (2026)
Morphing-based Compression for Data-centric ML Pipelines
par: Baunsgaard, Sebastian, et autres
Publié: (2025)
par: Baunsgaard, Sebastian, et autres
Publié: (2025)
Autonomous Electrochemistry Platform with Real-Time Normality Testing of Voltammetry Measurements Using ML
par: Al-Najjar, Anees, et autres
Publié: (2025)
par: Al-Najjar, Anees, et autres
Publié: (2025)
Ilargi: a GPU Compatible Factorized ML Model Training Framework
par: Sun, Wenbo, et autres
Publié: (2025)
par: Sun, Wenbo, et autres
Publié: (2025)
Characterization of GPU TEE Overheads in Distributed Data Parallel ML Training
par: Lee, Jonghyun, et autres
Publié: (2025)
par: Lee, Jonghyun, et autres
Publié: (2025)
Evolving HPC services to enable ML workloads on HPE Cray EX
par: Schuppli, Stefano, et autres
Publié: (2025)
par: Schuppli, Stefano, et autres
Publié: (2025)
Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis
par: Chu, Xiaoyu, et autres
Publié: (2024)
par: Chu, Xiaoyu, et autres
Publié: (2024)
Documents similaires
-
Compass: A Decentralized Scheduler for Latency-Sensitive ML Workflows
par: Yang, Yuting, et autres
Publié: (2024) -
Cloudless-Training: A Framework to Improve Efficiency of Geo-Distributed ML Training
par: Tan, Wenting, et autres
Publié: (2023) -
Scalable Runtime Architecture for Data-driven, Hybrid HPC and ML Workflow Applications
par: Merzky, Andre, et autres
Publié: (2025) -
HPAC-ML: A Programming Model for Embedding ML Surrogates in Scientific Applications
par: Fink, Zane, et autres
Publié: (2024) -
On-device Online Learning and Semantic Management of TinyML Systems
par: Ren, Haoyu, et autres
Publié: (2024)