FairBatching: Fairness-Aware Batch Formation for LLM Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lyu, Hongtao, Liu, Boyue, Wu, Mingyu, Chen, Haibo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching
von: Zheng, Zhen, et al.
Veröffentlicht: (2024)
von: Zheng, Zhen, et al.
Veröffentlicht: (2024)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
von: Zheng, Wanyi, et al.
Veröffentlicht: (2025)
von: Zheng, Wanyi, et al.
Veröffentlicht: (2025)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
von: Chen, Jiabin, et al.
Veröffentlicht: (2024)
von: Chen, Jiabin, et al.
Veröffentlicht: (2024)
Herring: Parallel Batch-Order-Fairness on DAG-based Blockchain Consensus
von: Putnik, Marko, et al.
Veröffentlicht: (2026)
von: Putnik, Marko, et al.
Veröffentlicht: (2026)
Constraint Programming Models For Serial Batch Scheduling With Minimum Batch Size
von: Huertas, Jorge A., et al.
Veröffentlicht: (2025)
von: Huertas, Jorge A., et al.
Veröffentlicht: (2025)
OrchMLLM: Orchestrate Multimodal Data with Batch Post-Balancing to Accelerate Multimodal Large Language Model Training
von: Zheng, Yijie, et al.
Veröffentlicht: (2025)
von: Zheng, Yijie, et al.
Veröffentlicht: (2025)
Equinox: Holistic Fair Scheduling in Serving Large Language Models
von: Wei, Zhixiang, et al.
Veröffentlicht: (2025)
von: Wei, Zhixiang, et al.
Veröffentlicht: (2025)
Token-Budget-Aware Pool Routing for Cost-Efficient LLM Inference
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
FairKV: Balancing Per-Head KV Cache for Fast Multi-GPU Inference
von: Zhao, Bingzhe, et al.
Veröffentlicht: (2025)
von: Zhao, Bingzhe, et al.
Veröffentlicht: (2025)
A Parallel CPU-GPU Framework for Batching Heuristic Operations in Depth-First Heuristic Search
von: Futuhi, Ehsan, et al.
Veröffentlicht: (2025)
von: Futuhi, Ehsan, et al.
Veröffentlicht: (2025)
On Using Large-Batches in Federated Learning
von: Tyagi, Sahil
Veröffentlicht: (2025)
von: Tyagi, Sahil
Veröffentlicht: (2025)
Multi-Agentic AI for Fairness-Aware and Accelerated Multi-modal Large Model Inference in Real-world Mobile Edge Networks
von: Li, Haiyuan, et al.
Veröffentlicht: (2026)
von: Li, Haiyuan, et al.
Veröffentlicht: (2026)
Justitia: Fair and Efficient Scheduling of Task-parallel LLM Agents with Selective Pampering
von: Yang, Mingyan, et al.
Veröffentlicht: (2025)
von: Yang, Mingyan, et al.
Veröffentlicht: (2025)
CONCUR: High-Throughput Agentic Batch Inference of LLM via Congestion-Based Concurrency Control
von: Chen, Qiaoling, et al.
Veröffentlicht: (2026)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2026)
Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching
von: Pang, Bowen, et al.
Veröffentlicht: (2025)
von: Pang, Bowen, et al.
Veröffentlicht: (2025)
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025)
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025)
Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation
von: Kim, Joon Ha, et al.
Veröffentlicht: (2026)
von: Kim, Joon Ha, et al.
Veröffentlicht: (2026)
AntBatchInfer: Elastic Batch Inference in the Kubernetes Cluster
von: Li, Siyuan, et al.
Veröffentlicht: (2024)
von: Li, Siyuan, et al.
Veröffentlicht: (2024)
ProMoE: Fast MoE-based LLM Serving using Proactive Caching
von: Song, Xiaoniu, et al.
Veröffentlicht: (2024)
von: Song, Xiaoniu, et al.
Veröffentlicht: (2024)
Fairness-Aware Job Scheduling for Multi-Job Federated Learning
von: Shi, Yuxin, et al.
Veröffentlicht: (2024)
von: Shi, Yuxin, et al.
Veröffentlicht: (2024)
DeServe: Towards Affordable Offline LLM Inference via Decentralization
von: Wu, Linyu, et al.
Veröffentlicht: (2025)
von: Wu, Linyu, et al.
Veröffentlicht: (2025)
Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures
von: Argerich, Mauricio Fadel, et al.
Veröffentlicht: (2026)
von: Argerich, Mauricio Fadel, et al.
Veröffentlicht: (2026)
Joint Optimization of Offloading, Batching and DVFS for Multiuser Co-Inference
von: Xu, Yaodan, et al.
Veröffentlicht: (2025)
von: Xu, Yaodan, et al.
Veröffentlicht: (2025)
KunServe: Parameter-centric Memory Management for Efficient Memory Overloading Handling in LLM Serving
von: Cheng, Rongxin, et al.
Veröffentlicht: (2024)
von: Cheng, Rongxin, et al.
Veröffentlicht: (2024)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
von: Liu, Xing, et al.
Veröffentlicht: (2025)
von: Liu, Xing, et al.
Veröffentlicht: (2025)
MineDraft: A Framework for Batch Parallel Speculative Decoding
von: Tang, Zhenwei, et al.
Veröffentlicht: (2026)
von: Tang, Zhenwei, et al.
Veröffentlicht: (2026)
Design a Win-Win Strategy That Is Fair to Both Service Providers and Tasks When Rejection Is Not an Option
von: Trabelsi, Yohai, et al.
Veröffentlicht: (2024)
von: Trabelsi, Yohai, et al.
Veröffentlicht: (2024)
GPU-Accelerated Batch-Dynamic Subgraph Matching
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
Hybrid Batch Normalisation: Resolving the Dilemma of Batch Normalisation in Federated Learning
von: Chen, Hongyao, et al.
Veröffentlicht: (2025)
von: Chen, Hongyao, et al.
Veröffentlicht: (2025)
LAPS: A Length-Aware-Prefill LLM Serving System
von: She, Jianshu, et al.
Veröffentlicht: (2026)
von: She, Jianshu, et al.
Veröffentlicht: (2026)
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
von: Li, Rongzhi, et al.
Veröffentlicht: (2025)
von: Li, Rongzhi, et al.
Veröffentlicht: (2025)
FedSAC: Dynamic Submodel Allocation for Collaborative Fairness in Federated Learning
von: Wang, Zihui, et al.
Veröffentlicht: (2024)
von: Wang, Zihui, et al.
Veröffentlicht: (2024)
Characterizing Mobile SoC for Accelerating Heterogeneous LLM Inference
von: Chen, Le, et al.
Veröffentlicht: (2025)
von: Chen, Le, et al.
Veröffentlicht: (2025)
Fast LLM Post-training via Decoupled and Fastest-of-N Speculation
von: Cheng, Rongxin, et al.
Veröffentlicht: (2025)
von: Cheng, Rongxin, et al.
Veröffentlicht: (2025)
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput
von: Song, Jingwei, et al.
Veröffentlicht: (2025)
von: Song, Jingwei, et al.
Veröffentlicht: (2025)
Accelerating LLM Inference with Precomputed Query Storage
von: Park, Jay H., et al.
Veröffentlicht: (2025)
von: Park, Jay H., et al.
Veröffentlicht: (2025)
Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference
von: Recasens, Pol G., et al.
Veröffentlicht: (2025)
von: Recasens, Pol G., et al.
Veröffentlicht: (2025)
FedFair^3: Unlocking Threefold Fairness in Federated Learning
von: Javaherian, Simin, et al.
Veröffentlicht: (2024)
von: Javaherian, Simin, et al.
Veröffentlicht: (2024)
GetBatch: Distributed Multi-Object Retrieval for ML Data Loading
von: Aizman, Alex, et al.
Veröffentlicht: (2026)
von: Aizman, Alex, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching
von: Zheng, Zhen, et al.
Veröffentlicht: (2024) -
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
von: Zheng, Wanyi, et al.
Veröffentlicht: (2025) -
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
von: Chen, Jiabin, et al.
Veröffentlicht: (2024) -
Herring: Parallel Batch-Order-Fairness on DAG-based Blockchain Consensus
von: Putnik, Marko, et al.
Veröffentlicht: (2026) -
Constraint Programming Models For Serial Batch Scheduling With Minimum Batch Size
von: Huertas, Jorge A., et al.
Veröffentlicht: (2025)