DARIS: An Oversubscribed Spatio-Temporal Scheduler for Real-Time DNN Inference on GPUs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Babaei, Amir Fakhim, Chantem, Thidapat |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SGPRS: Seamless GPU Partitioning Real-Time Scheduler for Periodic Deep Learning Workloads
von: Babaei, Amir Fakhim, et al.
Veröffentlicht: (2024)
von: Babaei, Amir Fakhim, et al.
Veröffentlicht: (2024)
ESG: Pipeline-Conscious Efficient Scheduling of DNN Workflows on Serverless Platforms with Shareable GPUs
von: Hui, Xinning, et al.
Veröffentlicht: (2024)
von: Hui, Xinning, et al.
Veröffentlicht: (2024)
Opara: Exploiting Operator Parallelism for Expediting DNN Inference on GPUs
von: Chen, Aodong, et al.
Veröffentlicht: (2023)
von: Chen, Aodong, et al.
Veröffentlicht: (2023)
Adaptive Heuristics for Scheduling DNN Inferencing on Edge and Cloud for Personalized UAV Fleets
von: Raj, Suman, et al.
Veröffentlicht: (2024)
von: Raj, Suman, et al.
Veröffentlicht: (2024)
Ocularone-Bench: Benchmarking DNN Models on GPUs to Assist the Visually Impaired
von: Raj, Suman, et al.
Veröffentlicht: (2025)
von: Raj, Suman, et al.
Veröffentlicht: (2025)
Preemption Aware Task Scheduling for Priority and Deadline Constrained DNN Inference Task Offloading in Homogeneous Mobile-Edge Networks
von: Cotter, Jamie, et al.
Veröffentlicht: (2025)
von: Cotter, Jamie, et al.
Veröffentlicht: (2025)
Serving Compound Inference Systems on Datacenter GPUs
von: Devata, Sriram, et al.
Veröffentlicht: (2026)
von: Devata, Sriram, et al.
Veröffentlicht: (2026)
FREESH: Fair, Resource- and Energy-Efficient Scheduling for LLM Serving on Heterogeneous GPUs
von: He, Xuan, et al.
Veröffentlicht: (2025)
von: He, Xuan, et al.
Veröffentlicht: (2025)
COHORT: Hybrid RL for Collaborative Large DNN Inference on Multi-Robot Systems Under Real-Time Constraints
von: Anwar, Mohammad Saeid, et al.
Veröffentlicht: (2026)
von: Anwar, Mohammad Saeid, et al.
Veröffentlicht: (2026)
Harpagon: Minimizing DNN Serving Cost via Efficient Dispatching, Scheduling and Splitting
von: Zhao, Zhixin, et al.
Veröffentlicht: (2024)
von: Zhao, Zhixin, et al.
Veröffentlicht: (2024)
SparOA: Sparse and Operator-aware Hybrid Scheduling for Edge DNN Inference
von: Zhang, Ziyang, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyang, et al.
Veröffentlicht: (2025)
An Online Fragmentation-Aware Scheduler for Managing GPU-Sharing Workloads on Multi-Instance GPUs
von: Ting, Hsu-Tzu, et al.
Veröffentlicht: (2025)
von: Ting, Hsu-Tzu, et al.
Veröffentlicht: (2025)
Fulcrum: Optimizing Concurrent DNN Training and Inferencing on Edge Accelerators
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)
Performance Characterization of Containerized DNN Training and Inference on Edge Accelerators
von: K., Prashanthi S., et al.
Veröffentlicht: (2023)
von: K., Prashanthi S., et al.
Veröffentlicht: (2023)
IsoSched: Preemptive Tile Cascaded Scheduling of Multi-DNN via Subgraph Isomorphism
von: Zhao, Boran, et al.
Veröffentlicht: (2025)
von: Zhao, Boran, et al.
Veröffentlicht: (2025)
A Real-Time Digital Twin for Adaptive Scheduling
von: Zhang, Yihe, et al.
Veröffentlicht: (2025)
von: Zhang, Yihe, et al.
Veröffentlicht: (2025)
Evaluating Multi-Instance DNN Inferencing on Multiple Accelerators of an Edge Device
von: Tayal, Mumuksh, et al.
Veröffentlicht: (2025)
von: Tayal, Mumuksh, et al.
Veröffentlicht: (2025)
Online Optimization of DNN Inference Network Utility in Collaborative Edge Computing
von: Li, Rui, et al.
Veröffentlicht: (2024)
von: Li, Rui, et al.
Veröffentlicht: (2024)
Collaborative Inference in DNN-based Satellite Systems with Dynamic Task Streams
von: Guan, Jinglong, et al.
Veröffentlicht: (2023)
von: Guan, Jinglong, et al.
Veröffentlicht: (2023)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
von: Chen, Jiabin, et al.
Veröffentlicht: (2024)
von: Chen, Jiabin, et al.
Veröffentlicht: (2024)
AdaOper: Energy-efficient and Responsive Concurrent DNN Inference on Mobile Devices
von: Lin, Zheng, et al.
Veröffentlicht: (2024)
von: Lin, Zheng, et al.
Veröffentlicht: (2024)
Where to Split? A Pareto-Front Analysis of DNN Partitioning for Edge Inference
von: Masud, Adiba, et al.
Veröffentlicht: (2026)
von: Masud, Adiba, et al.
Veröffentlicht: (2026)
APEX: Asynchronous Parallel CPU-GPU Execution for Online LLM Inference on Constrained GPUs
von: Fan, Jiakun, et al.
Veröffentlicht: (2025)
von: Fan, Jiakun, et al.
Veröffentlicht: (2025)
Infer-EDGE: Dynamic DNN Inference Optimization in 'Just-in-time' Edge-AI Implementations
von: Mounesan, Motahare, et al.
Veröffentlicht: (2025)
von: Mounesan, Motahare, et al.
Veröffentlicht: (2025)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
von: Chen, Xing, et al.
Veröffentlicht: (2025)
von: Chen, Xing, et al.
Veröffentlicht: (2025)
ParvaGPU: Efficient Spatial GPU Sharing for Large-Scale DNN Inference in Cloud Environments
von: Lee, Munkyu, et al.
Veröffentlicht: (2024)
von: Lee, Munkyu, et al.
Veröffentlicht: (2024)
Adaptive Device-Edge Collaboration on DNN Inference in AIoT: A Digital Twin-Assisted Approach
von: Hu, Shisheng, et al.
Veröffentlicht: (2024)
von: Hu, Shisheng, et al.
Veröffentlicht: (2024)
Pagoda: An Energy and Time Roofline Study for DNN Workloads on Edge Accelerators
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)
GCAPS: GPU Context-Aware Preemptive Priority-based Scheduling for Real-Time Tasks
von: Wang, Yidi, et al.
Veröffentlicht: (2024)
von: Wang, Yidi, et al.
Veröffentlicht: (2024)
Optimal Fixed Priority Scheduling in Multi-Stage Multi-Resource Distributed Real-Time Systems
von: Kumar, Niraj, et al.
Veröffentlicht: (2024)
von: Kumar, Niraj, et al.
Veröffentlicht: (2024)
SLO-Aware Scheduling for Large Language Model Inferences
von: Huang, Jinqi, et al.
Veröffentlicht: (2025)
von: Huang, Jinqi, et al.
Veröffentlicht: (2025)
An Adaptive Distributed Stencil Abstraction for GPUs
von: Bhosale, Aditya, et al.
Veröffentlicht: (2025)
von: Bhosale, Aditya, et al.
Veröffentlicht: (2025)
Accelerating Maximal Biclique Enumeration on GPUs
von: Hsieh, Chou-Ying, et al.
Veröffentlicht: (2024)
von: Hsieh, Chou-Ying, et al.
Veröffentlicht: (2024)
Parallelizing Maximal Clique Enumeration on GPUs
von: Almasri, Mohammad, et al.
Veröffentlicht: (2022)
von: Almasri, Mohammad, et al.
Veröffentlicht: (2022)
Optimizing sDTW for AMD GPUs
von: Latta-Lin, Daniel, et al.
Veröffentlicht: (2024)
von: Latta-Lin, Daniel, et al.
Veröffentlicht: (2024)
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
von: Wu, Yu, et al.
Veröffentlicht: (2025)
von: Wu, Yu, et al.
Veröffentlicht: (2025)
ReaLB: Real-Time Load Balancing for Multimodal MoE Inference
von: Wang, Yingping, et al.
Veröffentlicht: (2026)
von: Wang, Yingping, et al.
Veröffentlicht: (2026)
Practical Performance Guarantees for Pipelined DNN Inference
von: Archer, Aaron, et al.
Veröffentlicht: (2023)
von: Archer, Aaron, et al.
Veröffentlicht: (2023)
Symphony: Optimized DNN Model Serving using Deferred Batch Scheduling
von: Chen, Lequn, et al.
Veröffentlicht: (2023)
von: Chen, Lequn, et al.
Veröffentlicht: (2023)
Fast Kronecker Matrix-Matrix Multiplication on GPUs
von: Jangda, Abhinav, et al.
Veröffentlicht: (2024)
von: Jangda, Abhinav, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SGPRS: Seamless GPU Partitioning Real-Time Scheduler for Periodic Deep Learning Workloads
von: Babaei, Amir Fakhim, et al.
Veröffentlicht: (2024) -
ESG: Pipeline-Conscious Efficient Scheduling of DNN Workflows on Serverless Platforms with Shareable GPUs
von: Hui, Xinning, et al.
Veröffentlicht: (2024) -
Opara: Exploiting Operator Parallelism for Expediting DNN Inference on GPUs
von: Chen, Aodong, et al.
Veröffentlicht: (2023) -
Adaptive Heuristics for Scheduling DNN Inferencing on Edge and Cloud for Personalized UAV Fleets
von: Raj, Suman, et al.
Veröffentlicht: (2024) -
Ocularone-Bench: Benchmarking DNN Models on GPUs to Assist the Visually Impaired
von: Raj, Suman, et al.
Veröffentlicht: (2025)