Saved in:
| Main Authors: | Wen, Wei, Zhu, Quanyu, Chu, Weiwei, Chen, Wen-Yen, Yang, Jiyan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2409.04585 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Compass: A Decentralized Scheduler for Latency-Sensitive ML Workflows
by: Yang, Yuting, et al.
Published: (2024)
by: Yang, Yuting, et al.
Published: (2024)
Cloudless-Training: A Framework to Improve Efficiency of Geo-Distributed ML Training
by: Tan, Wenting, et al.
Published: (2023)
by: Tan, Wenting, et al.
Published: (2023)
Scalable Runtime Architecture for Data-driven, Hybrid HPC and ML Workflow Applications
by: Merzky, Andre, et al.
Published: (2025)
by: Merzky, Andre, et al.
Published: (2025)
On-device Online Learning and Semantic Management of TinyML Systems
by: Ren, Haoyu, et al.
Published: (2024)
by: Ren, Haoyu, et al.
Published: (2024)
Optimizing Split Learning Latency in TinyML-Based IoT Systems
by: Jenhani, Zied, et al.
Published: (2025)
by: Jenhani, Zied, et al.
Published: (2025)
TrainMover: An Interruption-Resilient Runtime for ML Training
by: Lao, ChonLam, et al.
Published: (2024)
by: Lao, ChonLam, et al.
Published: (2024)
HPAC-ML: A Programming Model for Embedding ML Surrogates in Scientific Applications
by: Fink, Zane, et al.
Published: (2024)
by: Fink, Zane, et al.
Published: (2024)
Declarative Data Pipeline for Large Scale ML Services
by: Yang, Yunzhao, et al.
Published: (2025)
by: Yang, Yunzhao, et al.
Published: (2025)
GetBatch: Distributed Multi-Object Retrieval for ML Data Loading
by: Aizman, Alex, et al.
Published: (2026)
by: Aizman, Alex, et al.
Published: (2026)
ML-based Modeling to Predict I/O Performance on Different Storage Sub-systems
by: Xu, Yiheng, et al.
Published: (2023)
by: Xu, Yiheng, et al.
Published: (2023)
ML-based Adaptive Prefetching and Data Placement for US HEP Systems
by: Karanam, Venkat Sai Suman Lamba, et al.
Published: (2025)
by: Karanam, Venkat Sai Suman Lamba, et al.
Published: (2025)
ML-ECS: A Collaborative Multimodal Learning Framework for Edge-Cloud Synergies
by: Liu, Yuze, et al.
Published: (2026)
by: Liu, Yuze, et al.
Published: (2026)
A Performance Analyzer for a Public Cloud's ML-Augmented VM Allocator
by: Bostandoost, Roozbeh, et al.
Published: (2025)
by: Bostandoost, Roozbeh, et al.
Published: (2025)
Loki: A System for Serving ML Inference Pipelines with Hardware and Accuracy Scaling
by: Ahmad, Sohaib, et al.
Published: (2024)
by: Ahmad, Sohaib, et al.
Published: (2024)
Evaluating Cross-Architecture Performance Modeling of Distributed ML Workloads Using StableHLO
by: Svedas, Jonas, et al.
Published: (2026)
by: Svedas, Jonas, et al.
Published: (2026)
PCCL: Photonic circuit-switched collective communication for distributed ML
by: Kumar, Abhishek Vijaya, et al.
Published: (2025)
by: Kumar, Abhishek Vijaya, et al.
Published: (2025)
NetSenseML: Network-Adaptive Compression for Efficient Distributed Machine Learning
by: Wang, Yisu, et al.
Published: (2025)
by: Wang, Yisu, et al.
Published: (2025)
PAL: A Variability-Aware Policy for Scheduling ML Workloads in GPU Clusters
by: Jain, Rutwik, et al.
Published: (2024)
by: Jain, Rutwik, et al.
Published: (2024)
OptiML: An End-to-End Framework for Program Synthesis and CUDA Kernel Optimization
by: Bhattacharjee, Arijit, et al.
Published: (2026)
by: Bhattacharjee, Arijit, et al.
Published: (2026)
Reimagining RDMA Through the Lens of ML
by: Warraich, Ertza, et al.
Published: (2025)
by: Warraich, Ertza, et al.
Published: (2025)
Deploying Graph Neural Networks in Wireless Networks: A Link Stability Viewpoint
by: Li, Jun, et al.
Published: (2024)
by: Li, Jun, et al.
Published: (2024)
An Efficient Heterogeneous Co-Design for Fine-Tuning on a Single GPU
by: Yang, Ruijia, et al.
Published: (2026)
by: Yang, Ruijia, et al.
Published: (2026)
PowerTrip: Exploiting Federated Heterogeneous Datacenter Power for Distributed ML Training
by: Mehboob, Talha, et al.
Published: (2025)
by: Mehboob, Talha, et al.
Published: (2025)
Flint: Compiler Enabled Cluster-Free Design Space Exploration for Distributed ML
by: Yoo, Jinsun, et al.
Published: (2026)
by: Yoo, Jinsun, et al.
Published: (2026)
Agentic TinyML for Intent-aware Handover in 6G Wireless Networks
by: Saleh, Alaa, et al.
Published: (2025)
by: Saleh, Alaa, et al.
Published: (2025)
Agentic Operator Generation for ML ASICs
by: Hammond, Alec M., et al.
Published: (2025)
by: Hammond, Alec M., et al.
Published: (2025)
Snowpark: Performant, Secure, User-Friendly Data Engineering and AI/ML Next To Your Data
by: Baker, Brandon, et al.
Published: (2025)
by: Baker, Brandon, et al.
Published: (2025)
tf.data service: A Case for Disaggregating ML Input Data Processing
by: Audibert, Andrew, et al.
Published: (2022)
by: Audibert, Andrew, et al.
Published: (2022)
yProv4ML: Effortless Provenance Tracking for Machine Learning Systems
by: Padovani, Gabriele, et al.
Published: (2025)
by: Padovani, Gabriele, et al.
Published: (2025)
Magneton: Optimizing Energy Efficiency of ML Systems via Differential Energy Debugging
by: Pan, Yi, et al.
Published: (2025)
by: Pan, Yi, et al.
Published: (2025)
Optimizing Distributed ML Communication with Fused Computation-Collective Operations
by: Punniyamurthy, Kishore, et al.
Published: (2023)
by: Punniyamurthy, Kishore, et al.
Published: (2023)
Taming the Memory Beast: Strategies for Reliable ML Training on Kubernetes
by: Ray, Jaideep
Published: (2024)
by: Ray, Jaideep
Published: (2024)
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
by: Agrawal, Anirudha, et al.
Published: (2024)
by: Agrawal, Anirudha, et al.
Published: (2024)
A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators
by: Colagrande, Luca, et al.
Published: (2026)
by: Colagrande, Luca, et al.
Published: (2026)
Morphing-based Compression for Data-centric ML Pipelines
by: Baunsgaard, Sebastian, et al.
Published: (2025)
by: Baunsgaard, Sebastian, et al.
Published: (2025)
A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via Faro
by: Jeon, Beomyeol, et al.
Published: (2024)
by: Jeon, Beomyeol, et al.
Published: (2024)
ML-QLS: Multilevel Quantum Layout Synthesis
by: Lin, Wan-Hsuan, et al.
Published: (2024)
by: Lin, Wan-Hsuan, et al.
Published: (2024)
Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis
by: Chu, Xiaoyu, et al.
Published: (2024)
by: Chu, Xiaoyu, et al.
Published: (2024)
SAIR: Cost-Efficient Multi-Stage ML Pipeline Autoscaling via In-Context Reinforcement Learning
by: Su, Jianchang, et al.
Published: (2026)
by: Su, Jianchang, et al.
Published: (2026)
Ilargi: a GPU Compatible Factorized ML Model Training Framework
by: Sun, Wenbo, et al.
Published: (2025)
by: Sun, Wenbo, et al.
Published: (2025)
Similar Items
-
Compass: A Decentralized Scheduler for Latency-Sensitive ML Workflows
by: Yang, Yuting, et al.
Published: (2024) -
Cloudless-Training: A Framework to Improve Efficiency of Geo-Distributed ML Training
by: Tan, Wenting, et al.
Published: (2023) -
Scalable Runtime Architecture for Data-driven, Hybrid HPC and ML Workflow Applications
by: Merzky, Andre, et al.
Published: (2025) -
On-device Online Learning and Semantic Management of TinyML Systems
by: Ren, Haoyu, et al.
Published: (2024) -
Optimizing Split Learning Latency in TinyML-Based IoT Systems
by: Jenhani, Zied, et al.
Published: (2025)