Multi-Bin Batching for Increasing LLM Inference Throughput
Fuente:
arXiv
Saved in:
| Main Authors: | Guldogan, Ozgur, Kunde, Jackson, Lee, Kangwook, Pedarsani, Ramtin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Quantized Decentralized Stochastic Learning over Directed Graphs
by: Taheri, Hossein, et al.
Published: (2020)
by: Taheri, Hossein, et al.
Published: (2020)
BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching
by: Zheng, Zhen, et al.
Published: (2024)
by: Zheng, Zhen, et al.
Published: (2024)
SMDP-Based Dynamic Batching for Improving Responsiveness and Energy Efficiency of Batch Services
by: Xu, Yaodan, et al.
Published: (2025)
by: Xu, Yaodan, et al.
Published: (2025)
Robust Decentralized Learning with Local Updates and Gradient Tracking
by: Ghiasvand, Sajjad, et al.
Published: (2024)
by: Ghiasvand, Sajjad, et al.
Published: (2024)
MoE-Gen: High-Throughput MoE Inference on a Single GPU with Module-Based Batching
by: Xu, Tairan, et al.
Published: (2025)
by: Xu, Tairan, et al.
Published: (2025)
Staggered Batch Scheduling: Co-optimizing Time-to-First-Token and Throughput for High-Efficiency LLM Inference
by: Tian, Jian, et al.
Published: (2025)
by: Tian, Jian, et al.
Published: (2025)
JITServe: SLO-aware LLM Serving with Imprecise Request Information
by: Zhang, Wei, et al.
Published: (2025)
by: Zhang, Wei, et al.
Published: (2025)
Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference
by: Recasens, Pol G., et al.
Published: (2025)
by: Recasens, Pol G., et al.
Published: (2025)
A Multi-Level Approach for Class Imbalance Problem in Federated Learning for Remote Industry 4.0 Applications
by: Hussain, Razin Farhan, et al.
Published: (2024)
by: Hussain, Razin Farhan, et al.
Published: (2024)
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
by: Agrawal, Amey, et al.
Published: (2024)
by: Agrawal, Amey, et al.
Published: (2024)
AntBatchInfer: Elastic Batch Inference in the Kubernetes Cluster
by: Li, Siyuan, et al.
Published: (2024)
by: Li, Siyuan, et al.
Published: (2024)
PackInfer: Compute- and I/O-Efficient Attention for Batched LLM Inference
by: Ning, Rui, et al.
Published: (2026)
by: Ning, Rui, et al.
Published: (2026)
PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference
by: Patel, Ishan, et al.
Published: (2026)
by: Patel, Ishan, et al.
Published: (2026)
LLM-Pilot: Characterize and Optimize Performance of your LLM Inference Services
by: Łazuka, Małgorzata, et al.
Published: (2024)
by: Łazuka, Małgorzata, et al.
Published: (2024)
FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
by: Oliaro, Gabriele, et al.
Published: (2024)
by: Oliaro, Gabriele, et al.
Published: (2024)
Technical Report on Reinforcement Learning Control on the Lucas-Nülle Inverted Pendulum
by: Schenke, Maximilian, et al.
Published: (2024)
by: Schenke, Maximilian, et al.
Published: (2024)
Measurement of Generative AI Workload Power Profiles for Whole-Facility Data Center Infrastructure Planning
by: Vercellino, Roberto, et al.
Published: (2026)
by: Vercellino, Roberto, et al.
Published: (2026)
Prediction of Permissioned Blockchain Performance for Resource Scaling Configurations
by: Jung, Seungwoo, et al.
Published: (2025)
by: Jung, Seungwoo, et al.
Published: (2025)
PipeInfer: Accelerating LLM Inference using Asynchronous Pipelined Speculation
by: Butler, Branden, et al.
Published: (2024)
by: Butler, Branden, et al.
Published: (2024)
SpecKV: Adaptive Speculative Decoding with Compression-Aware Gamma Selection
by: Shukla, Shikhar
Published: (2026)
by: Shukla, Shikhar
Published: (2026)
ScalableHD: Scalable and High-Throughput Hyperdimensional Computing Inference on Multi-Core CPUs
by: Parikh, Dhruv, et al.
Published: (2025)
by: Parikh, Dhruv, et al.
Published: (2025)
Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching
by: Pang, Bowen, et al.
Published: (2025)
by: Pang, Bowen, et al.
Published: (2025)
Towards Building the Federated GPT: Federated Instruction Tuning
by: Zhang, Jianyi, et al.
Published: (2023)
by: Zhang, Jianyi, et al.
Published: (2023)
OpenFedLLM: Training Large Language Models on Decentralized Private Data via Federated Learning
by: Ye, Rui, et al.
Published: (2024)
by: Ye, Rui, et al.
Published: (2024)
AMUSD: Asynchronous Multi-Device Speculative Decoding for LLM Acceleration
by: McDanel, Bradley
Published: (2024)
by: McDanel, Bradley
Published: (2024)
CONCUR: High-Throughput Agentic Batch Inference of LLM via Congestion-Based Concurrency Control
by: Chen, Qiaoling, et al.
Published: (2026)
by: Chen, Qiaoling, et al.
Published: (2026)
Multi-Source to Multi-Target Decentralized Federated Domain Adaptation
by: Wang, Su, et al.
Published: (2023)
by: Wang, Su, et al.
Published: (2023)
Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques
by: Behera, Adarsh Prasad, et al.
Published: (2025)
by: Behera, Adarsh Prasad, et al.
Published: (2025)
Scalable and Cost-Efficient ML Inference: Parallel Batch Processing with Serverless Functions
by: Barrak, Amine, et al.
Published: (2025)
by: Barrak, Amine, et al.
Published: (2025)
ISO: Overlap of Computation and Communication within Seqenence For LLM Inference
by: Xiao, Bin, et al.
Published: (2024)
by: Xiao, Bin, et al.
Published: (2024)
SPARe: Stacked Parallelism with Adaptive Reordering for Fault-Tolerant LLM Pretraining Systems with 100k+ GPUs
by: Lee, Jin, et al.
Published: (2026)
by: Lee, Jin, et al.
Published: (2026)
AMV-L: Lifecycle-Managed Agent Memory for Tail-Latency Control in Long-Running LLM Systems
by: Bamidele, Emmanuel
Published: (2026)
by: Bamidele, Emmanuel
Published: (2026)
Decentralized Online Learning for Random Inverse Problems Over Graphs
by: Zhang, Xiwei, et al.
Published: (2023)
by: Zhang, Xiwei, et al.
Published: (2023)
Token Management in Multi-Tenant AI Inference Platforms
by: Cunningham, William J.
Published: (2026)
by: Cunningham, William J.
Published: (2026)
LatencyPrism: Online Non-intrusive Latency Sculpting for SLO-Guaranteed LLM Inference
by: Du, Yin, et al.
Published: (2026)
by: Du, Yin, et al.
Published: (2026)
Deoxys: A Causal Inference Engine for Unhealthy Node Mitigation in Large-scale Cloud Infrastructure
by: Zhang, Chaoyun, et al.
Published: (2024)
by: Zhang, Chaoyun, et al.
Published: (2024)
Equilibrium in the Computing Continuum through Active Inference
by: Sedlak, Boris, et al.
Published: (2023)
by: Sedlak, Boris, et al.
Published: (2023)
Lynx: Enabling Efficient MoE Inference through Dynamic Batch-Aware Expert Selection
by: Gupta, Vima, et al.
Published: (2024)
by: Gupta, Vima, et al.
Published: (2024)
Understanding and Improving Communication Performance in Multi-node LLM Inference
by: Singhania, Prajwal, et al.
Published: (2025)
by: Singhania, Prajwal, et al.
Published: (2025)
Hybrid Batch Normalisation: Resolving the Dilemma of Batch Normalisation in Federated Learning
by: Chen, Hongyao, et al.
Published: (2025)
by: Chen, Hongyao, et al.
Published: (2025)
Similar Items
-
Quantized Decentralized Stochastic Learning over Directed Graphs
by: Taheri, Hossein, et al.
Published: (2020) -
BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching
by: Zheng, Zhen, et al.
Published: (2024) -
SMDP-Based Dynamic Batching for Improving Responsiveness and Energy Efficiency of Batch Services
by: Xu, Yaodan, et al.
Published: (2025) -
Robust Decentralized Learning with Local Updates and Gradient Tracking
by: Ghiasvand, Sajjad, et al.
Published: (2024) -
MoE-Gen: High-Throughput MoE Inference on a Single GPU with Module-Based Batching
by: Xu, Tairan, et al.
Published: (2025)