Staggered Batch Scheduling: Co-optimizing Time-to-First-Token and Throughput for High-Efficiency LLM Inference
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Tian, Jian, Li, Shuailong, Cao, Yang, Cui, Wenbo, Zhu, Minghan, Wu, Wenkang, Zhang, Jianming, Wang, Yanpeng, Xiao, Zhiwen, Hou, Zhenyu, Shen, Dou |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching
par: Pang, Bowen, et autres
Publié: (2025)
par: Pang, Bowen, et autres
Publié: (2025)
CONCUR: High-Throughput Agentic Batch Inference of LLM via Congestion-Based Concurrency Control
par: Chen, Qiaoling, et autres
Publié: (2026)
par: Chen, Qiaoling, et autres
Publié: (2026)
Joint Optimization of Offloading, Batching and DVFS for Multiuser Co-Inference
par: Xu, Yaodan, et autres
Publié: (2025)
par: Xu, Yaodan, et autres
Publié: (2025)
BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching
par: Zheng, Zhen, et autres
Publié: (2024)
par: Zheng, Zhen, et autres
Publié: (2024)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
par: Chen, Jiabin, et autres
Publié: (2024)
par: Chen, Jiabin, et autres
Publié: (2024)
Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving
par: Cheng, Ke, et autres
Publié: (2024)
par: Cheng, Ke, et autres
Publié: (2024)
Batch-Schedule-Execute: On Optimizing Concurrent Deterministic Scheduling for Blockchains (Extended Version)
par: Hay, Yaron, et autres
Publié: (2024)
par: Hay, Yaron, et autres
Publié: (2024)
ACE-GNN: Adaptive GNN Co-Inference with System-Aware Scheduling in Dynamic Edge Environments
par: Zhou, Ao, et autres
Publié: (2025)
par: Zhou, Ao, et autres
Publié: (2025)
MoE-Gen: High-Throughput MoE Inference on a Single GPU with Module-Based Batching
par: Xu, Tairan, et autres
Publié: (2025)
par: Xu, Tairan, et autres
Publié: (2025)
On the Efficiency of Dynamic Transaction Scheduling in Blockchain Sharding
par: Adhikari, Ramesh, et autres
Publié: (2025)
par: Adhikari, Ramesh, et autres
Publié: (2025)
Shift Parallelism: Low-Latency, High-Throughput LLM Inference for Dynamic Workloads
par: Hidayetoglu, Mert, et autres
Publié: (2025)
par: Hidayetoglu, Mert, et autres
Publié: (2025)
SLO-Aware Scheduling for Large Language Model Inferences
par: Huang, Jinqi, et autres
Publié: (2025)
par: Huang, Jinqi, et autres
Publié: (2025)
TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM Inference
par: Zhang, Hongbin, et autres
Publié: (2025)
par: Zhang, Hongbin, et autres
Publié: (2025)
FairBatching: Fairness-Aware Batch Formation for LLM Inference
par: Lyu, Hongtao, et autres
Publié: (2025)
par: Lyu, Hongtao, et autres
Publié: (2025)
AntBatchInfer: Elastic Batch Inference in the Kubernetes Cluster
par: Li, Siyuan, et autres
Publié: (2024)
par: Li, Siyuan, et autres
Publié: (2024)
Towards Energy Efficient Co-Scheduling in HPC
par: Zheng, Zhong, et autres
Publié: (2026)
par: Zheng, Zhong, et autres
Publié: (2026)
A HPC Co-Scheduler with Reinforcement Learning
par: Souza, Abel, et autres
Publié: (2024)
par: Souza, Abel, et autres
Publié: (2024)
COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training
par: Sakip, Akhmed, et autres
Publié: (2026)
par: Sakip, Akhmed, et autres
Publié: (2026)
Round-optimal $n$-Block Broadcast Schedules in Logarithmic Time
par: Träff, Jesper Larsson
Publié: (2023)
par: Träff, Jesper Larsson
Publié: (2023)
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
par: Wu, Yu, et autres
Publié: (2025)
par: Wu, Yu, et autres
Publié: (2025)
Multi-Bin Batching for Increasing LLM Inference Throughput
par: Guldogan, Ozgur, et autres
Publié: (2024)
par: Guldogan, Ozgur, et autres
Publié: (2024)
Argus: Token Aware Distributed LLM Inference Optimization
par: Wu, Panlong, et autres
Publié: (2025)
par: Wu, Panlong, et autres
Publié: (2025)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
par: Wang, Qipeng
Publié: (2026)
par: Wang, Qipeng
Publié: (2026)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
par: Chow, Will
Publié: (2025)
par: Chow, Will
Publié: (2025)
Towards Fast Setup and High Throughput of GPU Serverless Computing
par: Zhao, Han, et autres
Publié: (2024)
par: Zhao, Han, et autres
Publié: (2024)
Scaling Up Throughput-oriented LLM Inference Applications on Heterogeneous Opportunistic GPU Clusters with Pervasive Context Management
par: Phung, Thanh Son, et autres
Publié: (2025)
par: Phung, Thanh Son, et autres
Publié: (2025)
Research on fault diagnosis and root cause analysis based on full stack observability
par: Hou, Jian
Publié: (2025)
par: Hou, Jian
Publié: (2025)
Adaptive Heuristics for Scheduling DNN Inferencing on Edge and Cloud for Personalized UAV Fleets
par: Raj, Suman, et autres
Publié: (2024)
par: Raj, Suman, et autres
Publié: (2024)
SneakPeek: Data-Aware Model Selection and Scheduling for Inference Serving on the Edge
par: Wolfrath, Joel, et autres
Publié: (2025)
par: Wolfrath, Joel, et autres
Publié: (2025)
DARIS: An Oversubscribed Spatio-Temporal Scheduler for Real-Time DNN Inference on GPUs
par: Babaei, Amir Fakhim, et autres
Publié: (2025)
par: Babaei, Amir Fakhim, et autres
Publié: (2025)
Semantic Parallelism: Redefining Efficient MoE Inference via Model-Data Co-Scheduling
par: Li, Yan, et autres
Publié: (2025)
par: Li, Yan, et autres
Publié: (2025)
Symphony: Optimized DNN Model Serving using Deferred Batch Scheduling
par: Chen, Lequn, et autres
Publié: (2023)
par: Chen, Lequn, et autres
Publié: (2023)
AdaOper: Energy-efficient and Responsive Concurrent DNN Inference on Mobile Devices
par: Lin, Zheng, et autres
Publié: (2024)
par: Lin, Zheng, et autres
Publié: (2024)
CoCoI: Distributed Coded Inference System for Straggler Mitigation
par: Liu, Xing, et autres
Publié: (2025)
par: Liu, Xing, et autres
Publié: (2025)
KV Cache Compression for Inference Efficiency in LLMs: A Review
par: Liu, Yanyu, et autres
Publié: (2025)
par: Liu, Yanyu, et autres
Publié: (2025)
Designing Co-operation in Systems of Hierarchical, Multi-objective Schedulers for Stream Processing
par: Dangwal, Animesh, et autres
Publié: (2025)
par: Dangwal, Animesh, et autres
Publié: (2025)
TokenSim: Enabling Hardware and Software Exploration for Large Language Model Inference Systems
par: Wu, Feiyang, et autres
Publié: (2025)
par: Wu, Feiyang, et autres
Publié: (2025)
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
par: Wang, Weiye, et autres
Publié: (2026)
par: Wang, Weiye, et autres
Publié: (2026)
Orchestrating Joint Offloading and Scheduling for Low-Latency Edge SLAM
par: Zhang, Yao, et autres
Publié: (2025)
par: Zhang, Yao, et autres
Publié: (2025)
AdaBridge: Dynamic Data and Computation Reuse for Efficient Multi-task DNN Co-evolution in Edge Systems
par: Wang, Lehao, et autres
Publié: (2024)
par: Wang, Lehao, et autres
Publié: (2024)
Documents similaires
-
Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching
par: Pang, Bowen, et autres
Publié: (2025) -
CONCUR: High-Throughput Agentic Batch Inference of LLM via Congestion-Based Concurrency Control
par: Chen, Qiaoling, et autres
Publié: (2026) -
Joint Optimization of Offloading, Batching and DVFS for Multiuser Co-Inference
par: Xu, Yaodan, et autres
Publié: (2025) -
BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching
par: Zheng, Zhen, et autres
Publié: (2024) -
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
par: Chen, Jiabin, et autres
Publié: (2024)