SMART: A Surrogate Model for Predicting Application Runtime in Dragonfly Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Xin, Rizzini, Pietro Lodi, Medya, Sourav, Lan, Zhiling |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Study of Workload Interference with Intelligent Routing on Dragonfly
by: Kang, Yao, et al.
Published: (2024)
by: Kang, Yao, et al.
Published: (2024)
Q-adaptive: A Multi-Agent Reinforcement Learning Based Routing on Dragonfly Network
by: Kang, Yao, et al.
Published: (2024)
by: Kang, Yao, et al.
Published: (2024)
A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability
by: Liu, Ruitao, et al.
Published: (2026)
by: Liu, Ruitao, et al.
Published: (2026)
Interpretable Modeling of Deep Reinforcement Learning Driven Scheduling
by: Li, Boyang, et al.
Published: (2024)
by: Li, Boyang, et al.
Published: (2024)
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
by: Su, Zhaoyuan, et al.
Published: (2025)
by: Su, Zhaoyuan, et al.
Published: (2025)
A Robust Power Model Training Framework for Cloud Native Runtime Energy Metric Exporter
by: Choochotkaew, Sunyanan, et al.
Published: (2024)
by: Choochotkaew, Sunyanan, et al.
Published: (2024)
Machine-Learning-Driven Runtime Optimization of BLAS Level 3 on Modern Multi-Core Systems
by: Xia, Yufan, et al.
Published: (2024)
by: Xia, Yufan, et al.
Published: (2024)
A Machine Learning Approach Towards Runtime Optimisation of Matrix Multiplication
by: Xia, Yufan, et al.
Published: (2026)
by: Xia, Yufan, et al.
Published: (2026)
Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training
by: Lu, Yishun, et al.
Published: (2026)
by: Lu, Yishun, et al.
Published: (2026)
OpenG2G: A Simulation Platform for AI Datacenter-Grid Runtime Coordination
by: Chung, Jae-Won, et al.
Published: (2026)
by: Chung, Jae-Won, et al.
Published: (2026)
Maya: Optimizing Deep Learning Training Workloads using GPU Runtime Emulation
by: Yarlagadda, Srihas, et al.
Published: (2025)
by: Yarlagadda, Srihas, et al.
Published: (2025)
DeepCQ: General-Purpose Deep-Surrogate Framework for Lossy Compression Quality Prediction
by: Mumenin, Khondoker Mirazul, et al.
Published: (2025)
by: Mumenin, Khondoker Mirazul, et al.
Published: (2025)
QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration
by: Imani, HamidReza, et al.
Published: (2025)
by: Imani, HamidReza, et al.
Published: (2025)
EcoShift: Performance-Aware Power Management for Power-Constrained Heterogeneous Systems
by: Zheng, Zhong, et al.
Published: (2026)
by: Zheng, Zhong, et al.
Published: (2026)
A Universal Load Balancing Principle and Its Application to Large Language Model Serving
by: Chen, Zixi, et al.
Published: (2026)
by: Chen, Zixi, et al.
Published: (2026)
AERIS: Argonne Earth Systems Model for Reliable and Skillful Predictions
by: Hatanpää, Väinö, et al.
Published: (2025)
by: Hatanpää, Väinö, et al.
Published: (2025)
Prompt-Aware Scheduling for Low-Latency LLM Serving
by: Tao, Yiheng, et al.
Published: (2025)
by: Tao, Yiheng, et al.
Published: (2025)
Semi-decentralized Federated Time Series Prediction with Client Availability Budgets
by: Bao, Yunkai, et al.
Published: (2025)
by: Bao, Yunkai, et al.
Published: (2025)
Time-Series Learning for Proactive Fault Prediction in Distributed Systems with Deep Neural Structures
by: Wang, Yang, et al.
Published: (2025)
by: Wang, Yang, et al.
Published: (2025)
Union: An Automatic Workload Manager for Accelerating Network Simulation
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
Comprehensive Evaluation of GNN Training Systems: A Data Management Perspective
by: Yuan, Hao, et al.
Published: (2023)
by: Yuan, Hao, et al.
Published: (2023)
More for Less: Integrating Capability-Predominant and Capacity-Predominant Computing
by: Zheng, Zhong, et al.
Published: (2025)
by: Zheng, Zhong, et al.
Published: (2025)
Towards Energy Efficient Co-Scheduling in HPC
by: Zheng, Zhong, et al.
Published: (2026)
by: Zheng, Zhong, et al.
Published: (2026)
Adaptive Approach to Enhance Machine Learning Scheduling Algorithms During Runtime Using Reinforcement Learning in Metascheduling Applications
by: Alshaer, Samer, et al.
Published: (2025)
by: Alshaer, Samer, et al.
Published: (2025)
DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training
by: Hu, Tianhao, et al.
Published: (2026)
by: Hu, Tianhao, et al.
Published: (2026)
Large Language Model Aided QoS Prediction for Service Recommendation
by: Liu, Huiying, et al.
Published: (2024)
by: Liu, Huiying, et al.
Published: (2024)
Coordinated Power Management on Heterogeneous Systems
by: Zheng, Zhong, et al.
Published: (2025)
by: Zheng, Zhong, et al.
Published: (2025)
Cornserve: A Distributed Serving System for Any-to-Any Multimodal Models
by: Chung, Jae-Won, et al.
Published: (2026)
by: Chung, Jae-Won, et al.
Published: (2026)
Comprehensive Performance Modeling and System Design Insights for Foundation Models
by: Subramanian, Shashank, et al.
Published: (2024)
by: Subramanian, Shashank, et al.
Published: (2024)
Using Diffusion Models as Generative Replay in Continual Federated Learning -- What will Happen?
by: Mei, Yongsheng, et al.
Published: (2024)
by: Mei, Yongsheng, et al.
Published: (2024)
Mirage Persistent Kernel: A Compiler and Runtime for Mega-Kernelizing Tensor Programs
by: Cheng, Xinhao, et al.
Published: (2025)
by: Cheng, Xinhao, et al.
Published: (2025)
SkipPredict: When to Invest in Predictions for Scheduling
by: Shahout, Rana, et al.
Published: (2024)
by: Shahout, Rana, et al.
Published: (2024)
Privacy-Preserving Federated Learning with Consistency via Knowledge Distillation Using Conditional Generator
by: Luo, Kangyang, et al.
Published: (2024)
by: Luo, Kangyang, et al.
Published: (2024)
Predicting Temporal Aspects of Movement for Predictive Replication in Fog Environments
by: Balitzki, Emil, et al.
Published: (2023)
by: Balitzki, Emil, et al.
Published: (2023)
Federated Behavioural Planes: Explaining the Evolution of Client Behaviour in Federated Learning
by: Fenoglio, Dario, et al.
Published: (2024)
by: Fenoglio, Dario, et al.
Published: (2024)
Learning Semantics, Not Addresses: Runtime Neural Prefetching for Far Memory
by: Huang, Yutong, et al.
Published: (2025)
by: Huang, Yutong, et al.
Published: (2025)
RaMP: Runtime-Aware Megakernel Polymorphism for Mixture-of-Experts
by: Sharma, Vyom, et al.
Published: (2026)
by: Sharma, Vyom, et al.
Published: (2026)
EARL: Efficient Agentic Reinforcement Learning Systems for Large Language Models
by: Tan, Zheyue, et al.
Published: (2025)
by: Tan, Zheyue, et al.
Published: (2025)
Improving the End-to-End Efficiency of Offline Inference for Multi-LLM Applications Based on Sampling and Simulation
by: Fang, Jingzhi, et al.
Published: (2025)
by: Fang, Jingzhi, et al.
Published: (2025)
DeFRiS: Silo-Cooperative IoT Applications Scheduling via Decentralized Federated Reinforcement Learning
by: Wang, Zhiyu, et al.
Published: (2026)
by: Wang, Zhiyu, et al.
Published: (2026)
Similar Items
-
Study of Workload Interference with Intelligent Routing on Dragonfly
by: Kang, Yao, et al.
Published: (2024) -
Q-adaptive: A Multi-Agent Reinforcement Learning Based Routing on Dragonfly Network
by: Kang, Yao, et al.
Published: (2024) -
A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability
by: Liu, Ruitao, et al.
Published: (2026) -
Interpretable Modeling of Deep Reinforcement Learning Driven Scheduling
by: Li, Boyang, et al.
Published: (2024) -
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
by: Su, Zhaoyuan, et al.
Published: (2025)