Symphony: Optimized DNN Model Serving using Deferred Batch Scheduling
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Chen, Lequn, Deng, Weixin, Canumalla, Anirudh, Xin, Yu, Zhuo, Danyang, Philipose, Matthai, Krishnamurthy, Arvind |
|---|---|
| Format: | Preprint |
| Publié: |
2023
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Cortex: Workflow-Aware Resource Pooling and Scheduling for Agentic Serving
par: Pagonas, Nikos, et autres
Publié: (2025)
par: Pagonas, Nikos, et autres
Publié: (2025)
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
par: Chen, Siyuan, et autres
Publié: (2025)
par: Chen, Siyuan, et autres
Publié: (2025)
Harpagon: Minimizing DNN Serving Cost via Efficient Dispatching, Scheduling and Splitting
par: Zhao, Zhixin, et autres
Publié: (2024)
par: Zhao, Zhixin, et autres
Publié: (2024)
EdgeServing: Deadline-Aware Multi-DNN Serving at the Edge
par: Cao, Jiahe, et autres
Publié: (2026)
par: Cao, Jiahe, et autres
Publié: (2026)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
par: Chen, Jiabin, et autres
Publié: (2024)
par: Chen, Jiabin, et autres
Publié: (2024)
PolyServe: Efficient Multi-SLO Serving at Scale
par: Zhu, Kan, et autres
Publié: (2025)
par: Zhu, Kan, et autres
Publié: (2025)
GENSERVE: Efficient Co-Serving of Heterogeneous Diffusion Model Workloads
par: Ye, Fanjiang, et autres
Publié: (2026)
par: Ye, Fanjiang, et autres
Publié: (2026)
Batch-Schedule-Execute: On Optimizing Concurrent Deterministic Scheduling for Blockchains (Extended Version)
par: Hay, Yaron, et autres
Publié: (2024)
par: Hay, Yaron, et autres
Publié: (2024)
NanoFlow: Towards Optimal Large Language Model Serving Throughput
par: Zhu, Kan, et autres
Publié: (2024)
par: Zhu, Kan, et autres
Publié: (2024)
Enabling Large Batch Size Training for DNN Models Beyond the Memory Limit While Maintaining Performance
par: Piao, XinYu, et autres
Publié: (2021)
par: Piao, XinYu, et autres
Publié: (2021)
HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
par: Dong, Xianzhe, et autres
Publié: (2025)
par: Dong, Xianzhe, et autres
Publié: (2025)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
par: Ye, Zihao, et autres
Publié: (2025)
par: Ye, Zihao, et autres
Publié: (2025)
Enabling Efficient Batch Serving for LMaaS via Generation Length Prediction
par: Cheng, Ke, et autres
Publié: (2024)
par: Cheng, Ke, et autres
Publié: (2024)
ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production
par: Xiang, Yuxing, et autres
Publié: (2025)
par: Xiang, Yuxing, et autres
Publié: (2025)
AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System
par: Bai, Fengyao, et autres
Publié: (2026)
par: Bai, Fengyao, et autres
Publié: (2026)
Adaptive Heuristics for Scheduling DNN Inferencing on Edge and Cloud for Personalized UAV Fleets
par: Raj, Suman, et autres
Publié: (2024)
par: Raj, Suman, et autres
Publié: (2024)
DARIS: An Oversubscribed Spatio-Temporal Scheduler for Real-Time DNN Inference on GPUs
par: Babaei, Amir Fakhim, et autres
Publié: (2025)
par: Babaei, Amir Fakhim, et autres
Publié: (2025)
SneakPeek: Data-Aware Model Selection and Scheduling for Inference Serving on the Edge
par: Wolfrath, Joel, et autres
Publié: (2025)
par: Wolfrath, Joel, et autres
Publié: (2025)
ESG: Pipeline-Conscious Efficient Scheduling of DNN Workflows on Serverless Platforms with Shareable GPUs
par: Hui, Xinning, et autres
Publié: (2024)
par: Hui, Xinning, et autres
Publié: (2024)
IsoSched: Preemptive Tile Cascaded Scheduling of Multi-DNN via Subgraph Isomorphism
par: Zhao, Boran, et autres
Publié: (2025)
par: Zhao, Boran, et autres
Publié: (2025)
Past-Future Scheduler for LLM Serving under SLA Guarantees
par: Gong, Ruihao, et autres
Publié: (2025)
par: Gong, Ruihao, et autres
Publié: (2025)
PROSERVE: Unified Multi-Priority Request Scheduling for LLM Serving
par: Huang, Weizhe, et autres
Publié: (2025)
par: Huang, Weizhe, et autres
Publié: (2025)
A Predictive and Synergistic Two-Layer Scheduling Framework for LLM Serving
par: Zhang, Yue, et autres
Publié: (2025)
par: Zhang, Yue, et autres
Publié: (2025)
Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving
par: Cheng, Ke, et autres
Publié: (2024)
par: Cheng, Ke, et autres
Publié: (2024)
fabric-lib: RDMA Point-to-Point Communication for LLM Systems
par: Licker, Nandor, et autres
Publié: (2025)
par: Licker, Nandor, et autres
Publié: (2025)
Optimal Fixed Priority Scheduling in Multi-Stage Multi-Resource Distributed Real-Time Systems
par: Kumar, Niraj, et autres
Publié: (2024)
par: Kumar, Niraj, et autres
Publié: (2024)
MoLink: Distributed and Efficient Serving Framework for Large Models
par: Jin, Lewei, et autres
Publié: (2025)
par: Jin, Lewei, et autres
Publié: (2025)
FREESH: Fair, Resource- and Energy-Efficient Scheduling for LLM Serving on Heterogeneous GPUs
par: He, Xuan, et autres
Publié: (2025)
par: He, Xuan, et autres
Publié: (2025)
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
par: Zhong, Yinmin, et autres
Publié: (2024)
par: Zhong, Yinmin, et autres
Publié: (2024)
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start
par: Liu, Xueshen, et autres
Publié: (2026)
par: Liu, Xueshen, et autres
Publié: (2026)
Conveyor: Efficient Tool-aware LLM Serving with Tool Partial Execution
par: Xu, Yechen, et autres
Publié: (2024)
par: Xu, Yechen, et autres
Publié: (2024)
Fulcrum: Optimizing Concurrent DNN Training and Inferencing on Edge Accelerators
par: K., Prashanthi S., et autres
Publié: (2025)
par: K., Prashanthi S., et autres
Publié: (2025)
Parcae: Proactive, Liveput-Optimized DNN Training on Preemptible Instances
par: Duan, Jiangfei, et autres
Publié: (2024)
par: Duan, Jiangfei, et autres
Publié: (2024)
HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
par: Lou, Chiheng, et autres
Publié: (2025)
par: Lou, Chiheng, et autres
Publié: (2025)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
par: Zheng, Wanyi, et autres
Publié: (2025)
par: Zheng, Wanyi, et autres
Publié: (2025)
DeepServe: Serverless Large Language Model Serving at Scale
par: Hu, Junhao, et autres
Publié: (2025)
par: Hu, Junhao, et autres
Publié: (2025)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
par: Peng, You, et autres
Publié: (2026)
par: Peng, You, et autres
Publié: (2026)
CascadeInfer: Length-Aware Scheduling of LLM Serving with Low Latency and Load Balancing
par: Yuan, Yitao, et autres
Publié: (2025)
par: Yuan, Yitao, et autres
Publié: (2025)
Equinox: Holistic Fair Scheduling in Serving Large Language Models
par: Wei, Zhixiang, et autres
Publié: (2025)
par: Wei, Zhixiang, et autres
Publié: (2025)
Preemption Aware Task Scheduling for Priority and Deadline Constrained DNN Inference Task Offloading in Homogeneous Mobile-Edge Networks
par: Cotter, Jamie, et autres
Publié: (2025)
par: Cotter, Jamie, et autres
Publié: (2025)
Documents similaires
-
Cortex: Workflow-Aware Resource Pooling and Scheduling for Agentic Serving
par: Pagonas, Nikos, et autres
Publié: (2025) -
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
par: Chen, Siyuan, et autres
Publié: (2025) -
Harpagon: Minimizing DNN Serving Cost via Efficient Dispatching, Scheduling and Splitting
par: Zhao, Zhixin, et autres
Publié: (2024) -
EdgeServing: Deadline-Aware Multi-DNN Serving at the Edge
par: Cao, Jiahe, et autres
Publié: (2026) -
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
par: Chen, Jiabin, et autres
Publié: (2024)