ScaleDL: Towards Scalable and Efficient Runtime Prediction for Distributed Deep Learning Workloads
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Xiaokai, Huang, Shaoyuan, Li, Yuting, Wang, Xiaofei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sentinel: Scheduling Live Streams with Proactive Anomaly Detection in Crowdsourced Cloud-Edge Platforms
by: Li, Yuting, et al.
Published: (2025)
by: Li, Yuting, et al.
Published: (2025)
CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters
by: Huang, Shaoyuan, et al.
Published: (2026)
by: Huang, Shaoyuan, et al.
Published: (2026)
Maya: Optimizing Deep Learning Training Workloads using GPU Runtime Emulation
by: Yarlagadda, Srihas, et al.
Published: (2025)
by: Yarlagadda, Srihas, et al.
Published: (2025)
SparDL: Distributed Deep Learning Training with Efficient Sparse Communication
by: Zhao, Minjun, et al.
Published: (2023)
by: Zhao, Minjun, et al.
Published: (2023)
MetaEformer: Unveiling and Leveraging Meta-patterns for Complex and Dynamic Systems Load Forecasting
by: Huang, Shaoyuan, et al.
Published: (2025)
by: Huang, Shaoyuan, et al.
Published: (2025)
Prediction-Assisted Online Distributed Deep Learning Workload Scheduling in GPU Clusters
by: Luo, Ziyue, et al.
Published: (2025)
by: Luo, Ziyue, et al.
Published: (2025)
HRS: Hybrid Representation Framework with Scheduling Awareness for Time Series Forecasting in Crowdsourced Cloud-Edge Platforms
by: Zhang, Tiancheng, et al.
Published: (2025)
by: Zhang, Tiancheng, et al.
Published: (2025)
Cerberus: Multi-Agent Reasoning and Coverage-Guided Exploration for Static Detection of Runtime Errors
by: Dhulipala, Hridya, et al.
Published: (2025)
by: Dhulipala, Hridya, et al.
Published: (2025)
LearnedWMP: Workload Memory Prediction Using Distribution of Query Templates
by: Quader, Shaikh, et al.
Published: (2024)
by: Quader, Shaikh, et al.
Published: (2024)
GraVAC: Adaptive Compression for Communication-Efficient Distributed DL Training
by: Tyagi, Sahil, et al.
Published: (2023)
by: Tyagi, Sahil, et al.
Published: (2023)
Hybrid Learning and Optimization-Based Dynamic Scheduling for DL Workloads on Heterogeneous GPU Clusters
by: Dongare, Shruti, et al.
Published: (2025)
by: Dongare, Shruti, et al.
Published: (2025)
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
by: Su, Zhaoyuan, et al.
Published: (2025)
by: Su, Zhaoyuan, et al.
Published: (2025)
Toward Scalable Multirobot Control: Fast Policy Learning in Distributed MPC
by: Zhang, Xinglong, et al.
Published: (2024)
by: Zhang, Xinglong, et al.
Published: (2024)
MobilityDL: A Review of Deep Learning From Trajectory Data
by: Graser, Anita, et al.
Published: (2024)
by: Graser, Anita, et al.
Published: (2024)
Beyond Myopia: Learning from Positive and Unlabeled Data through Holistic Predictive Trends
by: Wang, Xinrui, et al.
Published: (2023)
by: Wang, Xinrui, et al.
Published: (2023)
Online Ensemble Transformer for Accurate Cloud Workload Forecasting in Predictive Auto-Scaling
by: Chen, Jiadong, et al.
Published: (2025)
by: Chen, Jiadong, et al.
Published: (2025)
LoCoDL: Communication-Efficient Distributed Learning with Local Training and Compression
by: Condat, Laurent, et al.
Published: (2024)
by: Condat, Laurent, et al.
Published: (2024)
Taming the Tail: NoI Topology Synthesis for Mixed DL Workloads on Chiplet-Based Accelerators
by: Shukla, Arnav, et al.
Published: (2025)
by: Shukla, Arnav, et al.
Published: (2025)
Scaling Vision Transformers: Evaluating DeepSpeed for Image-Centric Workloads
by: Trinh, Huy, et al.
Published: (2026)
by: Trinh, Huy, et al.
Published: (2026)
FedHQ: Hybrid Runtime Quantization for Federated Learning
by: Zheng, Zihao, et al.
Published: (2025)
by: Zheng, Zihao, et al.
Published: (2025)
Green AI: A Preliminary Empirical Study on Energy Consumption in DL Models Across Different Runtime Infrastructures
by: Alizadeh, Negar, et al.
Published: (2024)
by: Alizadeh, Negar, et al.
Published: (2024)
InstMeter: An Instruction-Level Method to Predict Energy and Latency of DL Model Inference on MCUs
by: Liu, Hao, et al.
Published: (2026)
by: Liu, Hao, et al.
Published: (2026)
RepDL: Bit-level Reproducible Deep Learning Training and Inference
by: Xie, Peichen, et al.
Published: (2025)
by: Xie, Peichen, et al.
Published: (2025)
DeepGate3: Towards Scalable Circuit Representation Learning
by: Shi, Zhengyuan, et al.
Published: (2024)
by: Shi, Zhengyuan, et al.
Published: (2024)
Learning Locally Interacting Discrete Dynamical Systems: Towards Data-Efficient and Scalable Prediction
by: Kang, Beomseok, et al.
Published: (2024)
by: Kang, Beomseok, et al.
Published: (2024)
DL2Fence: Integrating Deep Learning and Frame Fusion for Enhanced Detection and Localization of Refined Denial-of-Service in Large-Scale NoCs
by: Wang, Haoyu, et al.
Published: (2024)
by: Wang, Haoyu, et al.
Published: (2024)
Rethinking the Role of Dynamic Sparse Training for Scalable Deep Reinforcement Learning
by: Ma, Guozheng, et al.
Published: (2025)
by: Ma, Guozheng, et al.
Published: (2025)
Interference-Aware Edge Runtime Prediction with Conformal Matrix Completion
by: Huang, Tianshu, et al.
Published: (2025)
by: Huang, Tianshu, et al.
Published: (2025)
Embedding-Based Federated Learning with Runtime Governance for Iron Deficiency Prediction
by: Zhang, Fan, et al.
Published: (2026)
by: Zhang, Fan, et al.
Published: (2026)
MelissaDL x Breed: Towards Data-Efficient On-line Supervised Training of Multi-parametric Surrogates with Active Learning
by: Dymchenko, Sofya, et al.
Published: (2024)
by: Dymchenko, Sofya, et al.
Published: (2024)
Scalable and Precise Patch Robustness Certification for Deep Learning Models with Top-k Predictions
by: Zhou, Qilin, et al.
Published: (2025)
by: Zhou, Qilin, et al.
Published: (2025)
SPAQ-DL-SLAM: Towards Optimizing Deep Learning-based SLAM for Resource-Constrained Embedded Platforms
by: Pudasaini, Niraj, et al.
Published: (2024)
by: Pudasaini, Niraj, et al.
Published: (2024)
NGDB-Zoo: Towards Efficient and Scalable Neural Graph Databases Training
by: Xie, Zhongwei, et al.
Published: (2026)
by: Xie, Zhongwei, et al.
Published: (2026)
Principled Approximation Methods for Efficient and Scalable Deep Learning
by: Savarese, Pedro
Published: (2025)
by: Savarese, Pedro
Published: (2025)
Modeling and Predicting Transistor Aging under Workload Dependency using Machine Learning
by: Genssler, Paul R., et al.
Published: (2022)
by: Genssler, Paul R., et al.
Published: (2022)
Deep Learning-Enhanced Preconditioning for Efficient Conjugate Gradient Solvers in Large-Scale PDE Systems
by: Li, Rui, et al.
Published: (2024)
by: Li, Rui, et al.
Published: (2024)
Efficient and Scalable Deep Reinforcement Learning for Mean Field Control Games
by: Peng, Nianli, et al.
Published: (2024)
by: Peng, Nianli, et al.
Published: (2024)
Causal Spatio-Temporal Prediction: An Effective and Efficient Multi-Modal Approach
by: Huang, Yuting, et al.
Published: (2025)
by: Huang, Yuting, et al.
Published: (2025)
Debiased Machine Learning for Conformal Prediction of Counterfactual Outcomes Under Runtime Confounding
by: Barnatchez, Keith, et al.
Published: (2026)
by: Barnatchez, Keith, et al.
Published: (2026)
Semi-Supervised Learning with Balanced Deep Representation Distributions
by: Li, Changchun, et al.
Published: (2026)
by: Li, Changchun, et al.
Published: (2026)
Similar Items
-
Sentinel: Scheduling Live Streams with Proactive Anomaly Detection in Crowdsourced Cloud-Edge Platforms
by: Li, Yuting, et al.
Published: (2025) -
CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters
by: Huang, Shaoyuan, et al.
Published: (2026) -
Maya: Optimizing Deep Learning Training Workloads using GPU Runtime Emulation
by: Yarlagadda, Srihas, et al.
Published: (2025) -
SparDL: Distributed Deep Learning Training with Efficient Sparse Communication
by: Zhao, Minjun, et al.
Published: (2023) -
MetaEformer: Unveiling and Leveraging Meta-patterns for Complex and Dynamic Systems Load Forecasting
by: Huang, Shaoyuan, et al.
Published: (2025)