SMART: A Surrogate Model for Predicting Application Runtime in Dragonfly Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Xin, Rizzini, Pietro Lodi, Medya, Sourav, Lan, Zhiling |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Study of Workload Interference with Intelligent Routing on Dragonfly
von: Kang, Yao, et al.
Veröffentlicht: (2024)
von: Kang, Yao, et al.
Veröffentlicht: (2024)
Q-adaptive: A Multi-Agent Reinforcement Learning Based Routing on Dragonfly Network
von: Kang, Yao, et al.
Veröffentlicht: (2024)
von: Kang, Yao, et al.
Veröffentlicht: (2024)
A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability
von: Liu, Ruitao, et al.
Veröffentlicht: (2026)
von: Liu, Ruitao, et al.
Veröffentlicht: (2026)
Interpretable Modeling of Deep Reinforcement Learning Driven Scheduling
von: Li, Boyang, et al.
Veröffentlicht: (2024)
von: Li, Boyang, et al.
Veröffentlicht: (2024)
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
von: Su, Zhaoyuan, et al.
Veröffentlicht: (2025)
von: Su, Zhaoyuan, et al.
Veröffentlicht: (2025)
A Robust Power Model Training Framework for Cloud Native Runtime Energy Metric Exporter
von: Choochotkaew, Sunyanan, et al.
Veröffentlicht: (2024)
von: Choochotkaew, Sunyanan, et al.
Veröffentlicht: (2024)
Machine-Learning-Driven Runtime Optimization of BLAS Level 3 on Modern Multi-Core Systems
von: Xia, Yufan, et al.
Veröffentlicht: (2024)
von: Xia, Yufan, et al.
Veröffentlicht: (2024)
A Machine Learning Approach Towards Runtime Optimisation of Matrix Multiplication
von: Xia, Yufan, et al.
Veröffentlicht: (2026)
von: Xia, Yufan, et al.
Veröffentlicht: (2026)
Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training
von: Lu, Yishun, et al.
Veröffentlicht: (2026)
von: Lu, Yishun, et al.
Veröffentlicht: (2026)
OpenG2G: A Simulation Platform for AI Datacenter-Grid Runtime Coordination
von: Chung, Jae-Won, et al.
Veröffentlicht: (2026)
von: Chung, Jae-Won, et al.
Veröffentlicht: (2026)
Maya: Optimizing Deep Learning Training Workloads using GPU Runtime Emulation
von: Yarlagadda, Srihas, et al.
Veröffentlicht: (2025)
von: Yarlagadda, Srihas, et al.
Veröffentlicht: (2025)
DeepCQ: General-Purpose Deep-Surrogate Framework for Lossy Compression Quality Prediction
von: Mumenin, Khondoker Mirazul, et al.
Veröffentlicht: (2025)
von: Mumenin, Khondoker Mirazul, et al.
Veröffentlicht: (2025)
QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration
von: Imani, HamidReza, et al.
Veröffentlicht: (2025)
von: Imani, HamidReza, et al.
Veröffentlicht: (2025)
EcoShift: Performance-Aware Power Management for Power-Constrained Heterogeneous Systems
von: Zheng, Zhong, et al.
Veröffentlicht: (2026)
von: Zheng, Zhong, et al.
Veröffentlicht: (2026)
A Universal Load Balancing Principle and Its Application to Large Language Model Serving
von: Chen, Zixi, et al.
Veröffentlicht: (2026)
von: Chen, Zixi, et al.
Veröffentlicht: (2026)
AERIS: Argonne Earth Systems Model for Reliable and Skillful Predictions
von: Hatanpää, Väinö, et al.
Veröffentlicht: (2025)
von: Hatanpää, Väinö, et al.
Veröffentlicht: (2025)
Prompt-Aware Scheduling for Low-Latency LLM Serving
von: Tao, Yiheng, et al.
Veröffentlicht: (2025)
von: Tao, Yiheng, et al.
Veröffentlicht: (2025)
Semi-decentralized Federated Time Series Prediction with Client Availability Budgets
von: Bao, Yunkai, et al.
Veröffentlicht: (2025)
von: Bao, Yunkai, et al.
Veröffentlicht: (2025)
Time-Series Learning for Proactive Fault Prediction in Distributed Systems with Deep Neural Structures
von: Wang, Yang, et al.
Veröffentlicht: (2025)
von: Wang, Yang, et al.
Veröffentlicht: (2025)
Union: An Automatic Workload Manager for Accelerating Network Simulation
von: Wang, Xin, et al.
Veröffentlicht: (2024)
von: Wang, Xin, et al.
Veröffentlicht: (2024)
Comprehensive Evaluation of GNN Training Systems: A Data Management Perspective
von: Yuan, Hao, et al.
Veröffentlicht: (2023)
von: Yuan, Hao, et al.
Veröffentlicht: (2023)
More for Less: Integrating Capability-Predominant and Capacity-Predominant Computing
von: Zheng, Zhong, et al.
Veröffentlicht: (2025)
von: Zheng, Zhong, et al.
Veröffentlicht: (2025)
Towards Energy Efficient Co-Scheduling in HPC
von: Zheng, Zhong, et al.
Veröffentlicht: (2026)
von: Zheng, Zhong, et al.
Veröffentlicht: (2026)
Adaptive Approach to Enhance Machine Learning Scheduling Algorithms During Runtime Using Reinforcement Learning in Metascheduling Applications
von: Alshaer, Samer, et al.
Veröffentlicht: (2025)
von: Alshaer, Samer, et al.
Veröffentlicht: (2025)
DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training
von: Hu, Tianhao, et al.
Veröffentlicht: (2026)
von: Hu, Tianhao, et al.
Veröffentlicht: (2026)
Large Language Model Aided QoS Prediction for Service Recommendation
von: Liu, Huiying, et al.
Veröffentlicht: (2024)
von: Liu, Huiying, et al.
Veröffentlicht: (2024)
Coordinated Power Management on Heterogeneous Systems
von: Zheng, Zhong, et al.
Veröffentlicht: (2025)
von: Zheng, Zhong, et al.
Veröffentlicht: (2025)
Cornserve: A Distributed Serving System for Any-to-Any Multimodal Models
von: Chung, Jae-Won, et al.
Veröffentlicht: (2026)
von: Chung, Jae-Won, et al.
Veröffentlicht: (2026)
Comprehensive Performance Modeling and System Design Insights for Foundation Models
von: Subramanian, Shashank, et al.
Veröffentlicht: (2024)
von: Subramanian, Shashank, et al.
Veröffentlicht: (2024)
Using Diffusion Models as Generative Replay in Continual Federated Learning -- What will Happen?
von: Mei, Yongsheng, et al.
Veröffentlicht: (2024)
von: Mei, Yongsheng, et al.
Veröffentlicht: (2024)
Mirage Persistent Kernel: A Compiler and Runtime for Mega-Kernelizing Tensor Programs
von: Cheng, Xinhao, et al.
Veröffentlicht: (2025)
von: Cheng, Xinhao, et al.
Veröffentlicht: (2025)
SkipPredict: When to Invest in Predictions for Scheduling
von: Shahout, Rana, et al.
Veröffentlicht: (2024)
von: Shahout, Rana, et al.
Veröffentlicht: (2024)
Privacy-Preserving Federated Learning with Consistency via Knowledge Distillation Using Conditional Generator
von: Luo, Kangyang, et al.
Veröffentlicht: (2024)
von: Luo, Kangyang, et al.
Veröffentlicht: (2024)
Predicting Temporal Aspects of Movement for Predictive Replication in Fog Environments
von: Balitzki, Emil, et al.
Veröffentlicht: (2023)
von: Balitzki, Emil, et al.
Veröffentlicht: (2023)
Federated Behavioural Planes: Explaining the Evolution of Client Behaviour in Federated Learning
von: Fenoglio, Dario, et al.
Veröffentlicht: (2024)
von: Fenoglio, Dario, et al.
Veröffentlicht: (2024)
Learning Semantics, Not Addresses: Runtime Neural Prefetching for Far Memory
von: Huang, Yutong, et al.
Veröffentlicht: (2025)
von: Huang, Yutong, et al.
Veröffentlicht: (2025)
RaMP: Runtime-Aware Megakernel Polymorphism for Mixture-of-Experts
von: Sharma, Vyom, et al.
Veröffentlicht: (2026)
von: Sharma, Vyom, et al.
Veröffentlicht: (2026)
EARL: Efficient Agentic Reinforcement Learning Systems for Large Language Models
von: Tan, Zheyue, et al.
Veröffentlicht: (2025)
von: Tan, Zheyue, et al.
Veröffentlicht: (2025)
Improving the End-to-End Efficiency of Offline Inference for Multi-LLM Applications Based on Sampling and Simulation
von: Fang, Jingzhi, et al.
Veröffentlicht: (2025)
von: Fang, Jingzhi, et al.
Veröffentlicht: (2025)
DeFRiS: Silo-Cooperative IoT Applications Scheduling via Decentralized Federated Reinforcement Learning
von: Wang, Zhiyu, et al.
Veröffentlicht: (2026)
von: Wang, Zhiyu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Study of Workload Interference with Intelligent Routing on Dragonfly
von: Kang, Yao, et al.
Veröffentlicht: (2024) -
Q-adaptive: A Multi-Agent Reinforcement Learning Based Routing on Dragonfly Network
von: Kang, Yao, et al.
Veröffentlicht: (2024) -
A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability
von: Liu, Ruitao, et al.
Veröffentlicht: (2026) -
Interpretable Modeling of Deep Reinforcement Learning Driven Scheduling
von: Li, Boyang, et al.
Veröffentlicht: (2024) -
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
von: Su, Zhaoyuan, et al.
Veröffentlicht: (2025)