AntBatchInfer: Elastic Batch Inference in the Kubernetes Cluster
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Siyuan, Xiao, Youshao, Meng, Fanzhuang, Ju, Lin, Liang, Lei, Wang, Lin, Zhou, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AntDT: A Self-Adaptive Distributed Training Framework for Leader and Straggler Nodes
by: Xiao, Youshao, et al.
Published: (2024)
by: Xiao, Youshao, et al.
Published: (2024)
PackInfer: Compute- and I/O-Efficient Attention for Batched LLM Inference
by: Ning, Rui, et al.
Published: (2026)
by: Ning, Rui, et al.
Published: (2026)
G-Meta: Distributed Meta Learning in GPU Clusters for Large-Scale Recommender Systems
by: Xiao, Youshao, et al.
Published: (2024)
by: Xiao, Youshao, et al.
Published: (2024)
Lynx: Enabling Efficient MoE Inference through Dynamic Batch-Aware Expert Selection
by: Gupta, Vima, et al.
Published: (2024)
by: Gupta, Vima, et al.
Published: (2024)
Hybrid Batch Normalisation: Resolving the Dilemma of Batch Normalisation in Federated Learning
by: Chen, Hongyao, et al.
Published: (2025)
by: Chen, Hongyao, et al.
Published: (2025)
Staggered Batch Scheduling: Co-optimizing Time-to-First-Token and Throughput for High-Efficiency LLM Inference
by: Tian, Jian, et al.
Published: (2025)
by: Tian, Jian, et al.
Published: (2025)
Scalable and Cost-Efficient ML Inference: Parallel Batch Processing with Serverless Functions
by: Barrak, Amine, et al.
Published: (2025)
by: Barrak, Amine, et al.
Published: (2025)
Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference
by: Recasens, Pol G., et al.
Published: (2025)
by: Recasens, Pol G., et al.
Published: (2025)
The Streaming Batch Model for Efficient and Fault-Tolerant Heterogeneous Execution
by: Luan, Frank Sifei, et al.
Published: (2025)
by: Luan, Frank Sifei, et al.
Published: (2025)
BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching
by: Zheng, Zhen, et al.
Published: (2024)
by: Zheng, Zhen, et al.
Published: (2024)
SMDP-Based Dynamic Batching for Improving Responsiveness and Energy Efficiency of Batch Services
by: Xu, Yaodan, et al.
Published: (2025)
by: Xu, Yaodan, et al.
Published: (2025)
Collaborative Batch Size Optimization for Federated Learning
by: Geimer, Arno, et al.
Published: (2025)
by: Geimer, Arno, et al.
Published: (2025)
Optimal Batch Allocation for Wireless Federated Learning
by: Song, Jaeyoung, et al.
Published: (2024)
by: Song, Jaeyoung, et al.
Published: (2024)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
by: Chen, Jiabin, et al.
Published: (2024)
by: Chen, Jiabin, et al.
Published: (2024)
FLAMMABLE: A Multi-Model Federated Learning Framework with Multi-Model Engagement and Adaptive Batch Sizes
by: Lin, Shouxu, et al.
Published: (2025)
by: Lin, Shouxu, et al.
Published: (2025)
MoE-Gen: High-Throughput MoE Inference on a Single GPU with Module-Based Batching
by: Xu, Tairan, et al.
Published: (2025)
by: Xu, Tairan, et al.
Published: (2025)
Hybrid Dual-Batch and Cyclic Progressive Learning for Efficient Distributed Training
by: Lu, Kuan-Wei, et al.
Published: (2025)
by: Lu, Kuan-Wei, et al.
Published: (2025)
Symphony: Optimized DNN Model Serving using Deferred Batch Scheduling
by: Chen, Lequn, et al.
Published: (2023)
by: Chen, Lequn, et al.
Published: (2023)
BatchWeave: A Consistent Object-Store-Native Data Plane for Large Foundation Model Training
by: Sun, Ting, et al.
Published: (2026)
by: Sun, Ting, et al.
Published: (2026)
Design and Implementation of an Automated Disaster-recovery System for a Kubernetes Cluster Using LSTM
by: Kim, Ji-Beom, et al.
Published: (2024)
by: Kim, Ji-Beom, et al.
Published: (2024)
SURGE: SuperBatch Unified Resource-efficient GPU Encoding for Heterogeneous Partitioned Data
by: Kapadia, Shashank, et al.
Published: (2026)
by: Kapadia, Shashank, et al.
Published: (2026)
DynLP: Parallel Dynamic Batch Update for Label Propagation in Semi-Supervised Learning
by: Shovan, S M, et al.
Published: (2026)
by: Shovan, S M, et al.
Published: (2026)
DYNAMIX: RL-based Adaptive Batch Size Optimization in Distributed Machine Learning Systems
by: Dai, Yuanjun, et al.
Published: (2025)
by: Dai, Yuanjun, et al.
Published: (2025)
On Using Large-Batches in Federated Learning
by: Tyagi, Sahil
Published: (2025)
by: Tyagi, Sahil
Published: (2025)
Graph Neural Network Training Systems: A Performance Comparison of Full-Graph and Mini-Batch
by: Bajaj, Saurabh, et al.
Published: (2024)
by: Bajaj, Saurabh, et al.
Published: (2024)
Aryl: An Elastic Cluster Scheduler for Deep Learning
by: Li, Jiamin, et al.
Published: (2022)
by: Li, Jiamin, et al.
Published: (2022)
Joint Optimization of Offloading, Batching and DVFS for Multiuser Co-Inference
by: Xu, Yaodan, et al.
Published: (2025)
by: Xu, Yaodan, et al.
Published: (2025)
Enabling Large Batch Size Training for DNN Models Beyond the Memory Limit While Maintaining Performance
by: Piao, XinYu, et al.
Published: (2021)
by: Piao, XinYu, et al.
Published: (2021)
Multi-Bin Batching for Increasing LLM Inference Throughput
by: Guldogan, Ozgur, et al.
Published: (2024)
by: Guldogan, Ozgur, et al.
Published: (2024)
CDFGNN: a Systematic Design of Cache-based Distributed Full-Batch Graph Neural Network Training with Communication Reduction
by: Zhang, Shuai, et al.
Published: (2024)
by: Zhang, Shuai, et al.
Published: (2024)
FairBatching: Fairness-Aware Batch Formation for LLM Inference
by: Lyu, Hongtao, et al.
Published: (2025)
by: Lyu, Hongtao, et al.
Published: (2025)
Taming the Memory Beast: Strategies for Reliable ML Training on Kubernetes
by: Ray, Jaideep
Published: (2024)
by: Ray, Jaideep
Published: (2024)
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
by: Zhu, Ruidong, et al.
Published: (2025)
by: Zhu, Ruidong, et al.
Published: (2025)
Scaling Deep Learning Research with Kubernetes on the NRP Nautilus HyperCluster
by: Hurt, J. Alex, et al.
Published: (2024)
by: Hurt, J. Alex, et al.
Published: (2024)
Mitigating Temporal Blindness in Kubernetes Autoscaling: An Attention-Double-LSTM Framework
by: Shaikh, Faraz, et al.
Published: (2026)
by: Shaikh, Faraz, et al.
Published: (2026)
SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips
by: Yu, Jiahuan, et al.
Published: (2026)
by: Yu, Jiahuan, et al.
Published: (2026)
Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching
by: Pang, Bowen, et al.
Published: (2025)
by: Pang, Bowen, et al.
Published: (2025)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
by: Ye, Zihao, et al.
Published: (2025)
by: Ye, Zihao, et al.
Published: (2025)
DynaTrain: Fast Online Parallelism Switching for Elastic LLM Training
by: Wang, Yuanqing, et al.
Published: (2026)
by: Wang, Yuanqing, et al.
Published: (2026)
HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference
by: Zhong, Shuzhang, et al.
Published: (2025)
by: Zhong, Shuzhang, et al.
Published: (2025)
Similar Items
-
AntDT: A Self-Adaptive Distributed Training Framework for Leader and Straggler Nodes
by: Xiao, Youshao, et al.
Published: (2024) -
PackInfer: Compute- and I/O-Efficient Attention for Batched LLM Inference
by: Ning, Rui, et al.
Published: (2026) -
G-Meta: Distributed Meta Learning in GPU Clusters for Large-Scale Recommender Systems
by: Xiao, Youshao, et al.
Published: (2024) -
Lynx: Enabling Efficient MoE Inference through Dynamic Batch-Aware Expert Selection
by: Gupta, Vima, et al.
Published: (2024) -
Hybrid Batch Normalisation: Resolving the Dilemma of Batch Normalisation in Federated Learning
by: Chen, Hongyao, et al.
Published: (2025)