BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching
Fuente:
arXiv
Salvato in:
| Autori principali: | Zheng, Zhen, Ji, Xin, Fang, Taosong, Zhou, Fanghao, Liu, Chuanjie, Peng, Gang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching
di: Pang, Bowen, et al.
Pubblicazione: (2025)
di: Pang, Bowen, et al.
Pubblicazione: (2025)
CONCUR: High-Throughput Agentic Batch Inference of LLM via Congestion-Based Concurrency Control
di: Chen, Qiaoling, et al.
Pubblicazione: (2026)
di: Chen, Qiaoling, et al.
Pubblicazione: (2026)
FairBatching: Fairness-Aware Batch Formation for LLM Inference
di: Lyu, Hongtao, et al.
Pubblicazione: (2025)
di: Lyu, Hongtao, et al.
Pubblicazione: (2025)
Staggered Batch Scheduling: Co-optimizing Time-to-First-Token and Throughput for High-Efficiency LLM Inference
di: Tian, Jian, et al.
Pubblicazione: (2025)
di: Tian, Jian, et al.
Pubblicazione: (2025)
Multi-Bin Batching for Increasing LLM Inference Throughput
di: Guldogan, Ozgur, et al.
Pubblicazione: (2024)
di: Guldogan, Ozgur, et al.
Pubblicazione: (2024)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
di: Chen, Jiabin, et al.
Pubblicazione: (2024)
di: Chen, Jiabin, et al.
Pubblicazione: (2024)
AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System
di: Bai, Fengyao, et al.
Pubblicazione: (2026)
di: Bai, Fengyao, et al.
Pubblicazione: (2026)
Joint Optimization of Offloading, Batching and DVFS for Multiuser Co-Inference
di: Xu, Yaodan, et al.
Pubblicazione: (2025)
di: Xu, Yaodan, et al.
Pubblicazione: (2025)
AntBatchInfer: Elastic Batch Inference in the Kubernetes Cluster
di: Li, Siyuan, et al.
Pubblicazione: (2024)
di: Li, Siyuan, et al.
Pubblicazione: (2024)
Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference
di: Recasens, Pol G., et al.
Pubblicazione: (2025)
di: Recasens, Pol G., et al.
Pubblicazione: (2025)
PackInfer: Compute- and I/O-Efficient Attention for Batched LLM Inference
di: Ning, Rui, et al.
Pubblicazione: (2026)
di: Ning, Rui, et al.
Pubblicazione: (2026)
Batch Query Processing and Optimization for Agentic Workflows
di: Shen, Junyi, et al.
Pubblicazione: (2025)
di: Shen, Junyi, et al.
Pubblicazione: (2025)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
di: Zheng, Wanyi, et al.
Pubblicazione: (2025)
di: Zheng, Wanyi, et al.
Pubblicazione: (2025)
GPU-Accelerated Batch-Dynamic Subgraph Matching
di: Qiu, Linshan, et al.
Pubblicazione: (2024)
di: Qiu, Linshan, et al.
Pubblicazione: (2024)
Hybrid Batch Normalisation: Resolving the Dilemma of Batch Normalisation in Federated Learning
di: Chen, Hongyao, et al.
Pubblicazione: (2025)
di: Chen, Hongyao, et al.
Pubblicazione: (2025)
Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs
di: Ekelund, Jonah, et al.
Pubblicazione: (2025)
di: Ekelund, Jonah, et al.
Pubblicazione: (2025)
Tangram: High-resolution Video Analytics on Serverless Platform with SLO-aware Batching
di: Peng, Haosong, et al.
Pubblicazione: (2024)
di: Peng, Haosong, et al.
Pubblicazione: (2024)
BSODiag: A Global Diagnosis Framework for Batch Servers Outage in Large-scale Cloud Infrastructure Systems
di: Duan, Tao, et al.
Pubblicazione: (2025)
di: Duan, Tao, et al.
Pubblicazione: (2025)
Enabling Efficient Batch Serving for LMaaS via Generation Length Prediction
di: Cheng, Ke, et al.
Pubblicazione: (2024)
di: Cheng, Ke, et al.
Pubblicazione: (2024)
Collaborative Batch Size Optimization for Federated Learning
di: Geimer, Arno, et al.
Pubblicazione: (2025)
di: Geimer, Arno, et al.
Pubblicazione: (2025)
MoE-Gen: High-Throughput MoE Inference on a Single GPU with Module-Based Batching
di: Xu, Tairan, et al.
Pubblicazione: (2025)
di: Xu, Tairan, et al.
Pubblicazione: (2025)
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference
di: Zhao, Yihao, et al.
Pubblicazione: (2025)
di: Zhao, Yihao, et al.
Pubblicazione: (2025)
Are Your Epochs Too Epic? Batch Free Can Be Harmful
di: Kim, Daewoo, et al.
Pubblicazione: (2024)
di: Kim, Daewoo, et al.
Pubblicazione: (2024)
Batched DGEMMs for scientific codes running on long vector architectures
di: Banchelli, Fabio, et al.
Pubblicazione: (2025)
di: Banchelli, Fabio, et al.
Pubblicazione: (2025)
Batch Denoising for AIGC Service Provisioning in Wireless Edge Networks
di: Xu, Jinghang, et al.
Pubblicazione: (2025)
di: Xu, Jinghang, et al.
Pubblicazione: (2025)
COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training
di: Sakip, Akhmed, et al.
Pubblicazione: (2026)
di: Sakip, Akhmed, et al.
Pubblicazione: (2026)
Schedule-Level Shared-Prefix Reuse for LLM RL Training
di: Li, Pengbo, et al.
Pubblicazione: (2026)
di: Li, Pengbo, et al.
Pubblicazione: (2026)
Symphony: Optimized DNN Model Serving using Deferred Batch Scheduling
di: Chen, Lequn, et al.
Pubblicazione: (2023)
di: Chen, Lequn, et al.
Pubblicazione: (2023)
A Reinforcement Learning Based Backfilling Strategy for HPC Batch Jobs
di: Kolker-Hicks, Elliot, et al.
Pubblicazione: (2024)
di: Kolker-Hicks, Elliot, et al.
Pubblicazione: (2024)
Herring: Parallel Batch-Order-Fairness on DAG-based Blockchain Consensus
di: Putnik, Marko, et al.
Pubblicazione: (2026)
di: Putnik, Marko, et al.
Pubblicazione: (2026)
Ark: Offchain Transaction Batching in Bitcoin
di: Keer, Pim, et al.
Pubblicazione: (2026)
di: Keer, Pim, et al.
Pubblicazione: (2026)
Batch-Schedule-Execute: On Optimizing Concurrent Deterministic Scheduling for Blockchains (Extended Version)
di: Hay, Yaron, et al.
Pubblicazione: (2024)
di: Hay, Yaron, et al.
Pubblicazione: (2024)
Constraint Programming Models For Serial Batch Scheduling With Minimum Batch Size
di: Huertas, Jorge A., et al.
Pubblicazione: (2025)
di: Huertas, Jorge A., et al.
Pubblicazione: (2025)
On Using Large-Batches in Federated Learning
di: Tyagi, Sahil
Pubblicazione: (2025)
di: Tyagi, Sahil
Pubblicazione: (2025)
Scalable and Cost-Efficient ML Inference: Parallel Batch Processing with Serverless Functions
di: Barrak, Amine, et al.
Pubblicazione: (2025)
di: Barrak, Amine, et al.
Pubblicazione: (2025)
Optimal Batch Allocation for Wireless Federated Learning
di: Song, Jaeyoung, et al.
Pubblicazione: (2024)
di: Song, Jaeyoung, et al.
Pubblicazione: (2024)
Argus: Token Aware Distributed LLM Inference Optimization
di: Wu, Panlong, et al.
Pubblicazione: (2025)
di: Wu, Panlong, et al.
Pubblicazione: (2025)
SMDP-Based Dynamic Batching for Improving Responsiveness and Energy Efficiency of Batch Services
di: Xu, Yaodan, et al.
Pubblicazione: (2025)
di: Xu, Yaodan, et al.
Pubblicazione: (2025)
Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches
di: Fang, Shaoke, et al.
Pubblicazione: (2026)
di: Fang, Shaoke, et al.
Pubblicazione: (2026)
On Optimal Batch Size in Coded Computing
di: Saha, Swapnil, et al.
Pubblicazione: (2025)
di: Saha, Swapnil, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching
di: Pang, Bowen, et al.
Pubblicazione: (2025) -
CONCUR: High-Throughput Agentic Batch Inference of LLM via Congestion-Based Concurrency Control
di: Chen, Qiaoling, et al.
Pubblicazione: (2026) -
FairBatching: Fairness-Aware Batch Formation for LLM Inference
di: Lyu, Hongtao, et al.
Pubblicazione: (2025) -
Staggered Batch Scheduling: Co-optimizing Time-to-First-Token and Throughput for High-Efficiency LLM Inference
di: Tian, Jian, et al.
Pubblicazione: (2025) -
Multi-Bin Batching for Increasing LLM Inference Throughput
di: Guldogan, Ozgur, et al.
Pubblicazione: (2024)