Multi-Bin Batching for Increasing LLM Inference Throughput
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guldogan, Ozgur, Kunde, Jackson, Lee, Kangwook, Pedarsani, Ramtin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Quantized Decentralized Stochastic Learning over Directed Graphs
von: Taheri, Hossein, et al.
Veröffentlicht: (2020)
von: Taheri, Hossein, et al.
Veröffentlicht: (2020)
BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching
von: Zheng, Zhen, et al.
Veröffentlicht: (2024)
von: Zheng, Zhen, et al.
Veröffentlicht: (2024)
SMDP-Based Dynamic Batching for Improving Responsiveness and Energy Efficiency of Batch Services
von: Xu, Yaodan, et al.
Veröffentlicht: (2025)
von: Xu, Yaodan, et al.
Veröffentlicht: (2025)
Robust Decentralized Learning with Local Updates and Gradient Tracking
von: Ghiasvand, Sajjad, et al.
Veröffentlicht: (2024)
von: Ghiasvand, Sajjad, et al.
Veröffentlicht: (2024)
MoE-Gen: High-Throughput MoE Inference on a Single GPU with Module-Based Batching
von: Xu, Tairan, et al.
Veröffentlicht: (2025)
von: Xu, Tairan, et al.
Veröffentlicht: (2025)
Staggered Batch Scheduling: Co-optimizing Time-to-First-Token and Throughput for High-Efficiency LLM Inference
von: Tian, Jian, et al.
Veröffentlicht: (2025)
von: Tian, Jian, et al.
Veröffentlicht: (2025)
JITServe: SLO-aware LLM Serving with Imprecise Request Information
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference
von: Recasens, Pol G., et al.
Veröffentlicht: (2025)
von: Recasens, Pol G., et al.
Veröffentlicht: (2025)
A Multi-Level Approach for Class Imbalance Problem in Federated Learning for Remote Industry 4.0 Applications
von: Hussain, Razin Farhan, et al.
Veröffentlicht: (2024)
von: Hussain, Razin Farhan, et al.
Veröffentlicht: (2024)
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
AntBatchInfer: Elastic Batch Inference in the Kubernetes Cluster
von: Li, Siyuan, et al.
Veröffentlicht: (2024)
von: Li, Siyuan, et al.
Veröffentlicht: (2024)
PackInfer: Compute- and I/O-Efficient Attention for Batched LLM Inference
von: Ning, Rui, et al.
Veröffentlicht: (2026)
von: Ning, Rui, et al.
Veröffentlicht: (2026)
PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference
von: Patel, Ishan, et al.
Veröffentlicht: (2026)
von: Patel, Ishan, et al.
Veröffentlicht: (2026)
LLM-Pilot: Characterize and Optimize Performance of your LLM Inference Services
von: Łazuka, Małgorzata, et al.
Veröffentlicht: (2024)
von: Łazuka, Małgorzata, et al.
Veröffentlicht: (2024)
FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2024)
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2024)
Technical Report on Reinforcement Learning Control on the Lucas-Nülle Inverted Pendulum
von: Schenke, Maximilian, et al.
Veröffentlicht: (2024)
von: Schenke, Maximilian, et al.
Veröffentlicht: (2024)
Measurement of Generative AI Workload Power Profiles for Whole-Facility Data Center Infrastructure Planning
von: Vercellino, Roberto, et al.
Veröffentlicht: (2026)
von: Vercellino, Roberto, et al.
Veröffentlicht: (2026)
Prediction of Permissioned Blockchain Performance for Resource Scaling Configurations
von: Jung, Seungwoo, et al.
Veröffentlicht: (2025)
von: Jung, Seungwoo, et al.
Veröffentlicht: (2025)
SpecKV: Adaptive Speculative Decoding with Compression-Aware Gamma Selection
von: Shukla, Shikhar
Veröffentlicht: (2026)
von: Shukla, Shikhar
Veröffentlicht: (2026)
PipeInfer: Accelerating LLM Inference using Asynchronous Pipelined Speculation
von: Butler, Branden, et al.
Veröffentlicht: (2024)
von: Butler, Branden, et al.
Veröffentlicht: (2024)
ScalableHD: Scalable and High-Throughput Hyperdimensional Computing Inference on Multi-Core CPUs
von: Parikh, Dhruv, et al.
Veröffentlicht: (2025)
von: Parikh, Dhruv, et al.
Veröffentlicht: (2025)
Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching
von: Pang, Bowen, et al.
Veröffentlicht: (2025)
von: Pang, Bowen, et al.
Veröffentlicht: (2025)
Towards Building the Federated GPT: Federated Instruction Tuning
von: Zhang, Jianyi, et al.
Veröffentlicht: (2023)
von: Zhang, Jianyi, et al.
Veröffentlicht: (2023)
OpenFedLLM: Training Large Language Models on Decentralized Private Data via Federated Learning
von: Ye, Rui, et al.
Veröffentlicht: (2024)
von: Ye, Rui, et al.
Veröffentlicht: (2024)
AMUSD: Asynchronous Multi-Device Speculative Decoding for LLM Acceleration
von: McDanel, Bradley
Veröffentlicht: (2024)
von: McDanel, Bradley
Veröffentlicht: (2024)
CONCUR: High-Throughput Agentic Batch Inference of LLM via Congestion-Based Concurrency Control
von: Chen, Qiaoling, et al.
Veröffentlicht: (2026)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2026)
Multi-Source to Multi-Target Decentralized Federated Domain Adaptation
von: Wang, Su, et al.
Veröffentlicht: (2023)
von: Wang, Su, et al.
Veröffentlicht: (2023)
Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques
von: Behera, Adarsh Prasad, et al.
Veröffentlicht: (2025)
von: Behera, Adarsh Prasad, et al.
Veröffentlicht: (2025)
Scalable and Cost-Efficient ML Inference: Parallel Batch Processing with Serverless Functions
von: Barrak, Amine, et al.
Veröffentlicht: (2025)
von: Barrak, Amine, et al.
Veröffentlicht: (2025)
ISO: Overlap of Computation and Communication within Seqenence For LLM Inference
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
AMV-L: Lifecycle-Managed Agent Memory for Tail-Latency Control in Long-Running LLM Systems
von: Bamidele, Emmanuel
Veröffentlicht: (2026)
von: Bamidele, Emmanuel
Veröffentlicht: (2026)
SPARe: Stacked Parallelism with Adaptive Reordering for Fault-Tolerant LLM Pretraining Systems with 100k+ GPUs
von: Lee, Jin, et al.
Veröffentlicht: (2026)
von: Lee, Jin, et al.
Veröffentlicht: (2026)
Decentralized Online Learning for Random Inverse Problems Over Graphs
von: Zhang, Xiwei, et al.
Veröffentlicht: (2023)
von: Zhang, Xiwei, et al.
Veröffentlicht: (2023)
Token Management in Multi-Tenant AI Inference Platforms
von: Cunningham, William J.
Veröffentlicht: (2026)
von: Cunningham, William J.
Veröffentlicht: (2026)
LatencyPrism: Online Non-intrusive Latency Sculpting for SLO-Guaranteed LLM Inference
von: Du, Yin, et al.
Veröffentlicht: (2026)
von: Du, Yin, et al.
Veröffentlicht: (2026)
Deoxys: A Causal Inference Engine for Unhealthy Node Mitigation in Large-scale Cloud Infrastructure
von: Zhang, Chaoyun, et al.
Veröffentlicht: (2024)
von: Zhang, Chaoyun, et al.
Veröffentlicht: (2024)
Equilibrium in the Computing Continuum through Active Inference
von: Sedlak, Boris, et al.
Veröffentlicht: (2023)
von: Sedlak, Boris, et al.
Veröffentlicht: (2023)
Lynx: Enabling Efficient MoE Inference through Dynamic Batch-Aware Expert Selection
von: Gupta, Vima, et al.
Veröffentlicht: (2024)
von: Gupta, Vima, et al.
Veröffentlicht: (2024)
Understanding and Improving Communication Performance in Multi-node LLM Inference
von: Singhania, Prajwal, et al.
Veröffentlicht: (2025)
von: Singhania, Prajwal, et al.
Veröffentlicht: (2025)
Hybrid Batch Normalisation: Resolving the Dilemma of Batch Normalisation in Federated Learning
von: Chen, Hongyao, et al.
Veröffentlicht: (2025)
von: Chen, Hongyao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Quantized Decentralized Stochastic Learning over Directed Graphs
von: Taheri, Hossein, et al.
Veröffentlicht: (2020) -
BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching
von: Zheng, Zhen, et al.
Veröffentlicht: (2024) -
SMDP-Based Dynamic Batching for Improving Responsiveness and Energy Efficiency of Batch Services
von: Xu, Yaodan, et al.
Veröffentlicht: (2025) -
Robust Decentralized Learning with Local Updates and Gradient Tracking
von: Ghiasvand, Sajjad, et al.
Veröffentlicht: (2024) -
MoE-Gen: High-Throughput MoE Inference on a Single GPU with Module-Based Batching
von: Xu, Tairan, et al.
Veröffentlicht: (2025)