Requests of a Feather Must Flock Together: Batch Size vs. Prefix Homogeneity in LLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Rathi, Saksham, Preeti, Vutukuru, Mythili |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reliability Estimation of News Media Sources: Birds of a Feather Flock Together
by: Burdisso, Sergio, et al.
Published: (2024)
by: Burdisso, Sergio, et al.
Published: (2024)
BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching
by: Zheng, Zhen, et al.
Published: (2024)
by: Zheng, Zhen, et al.
Published: (2024)
Hydragen: High-Throughput LLM Inference with Shared Prefixes
by: Juravsky, Jordan, et al.
Published: (2024)
by: Juravsky, Jordan, et al.
Published: (2024)
Full-Graph vs. Mini-Batch Training: Comprehensive Analysis from a Batch Size and Fan-Out Size Perspective
by: Liu, Mengfan, et al.
Published: (2026)
by: Liu, Mengfan, et al.
Published: (2026)
SCORPIO: Serving the Right Requests at the Right Time for Heterogeneous SLOs in LLM Inference
by: Tang, Yinghao, et al.
Published: (2025)
by: Tang, Yinghao, et al.
Published: (2025)
RW-TTT: Batched Serving for Request-Owned Test-Time Training State
by: Yang, Jian, et al.
Published: (2026)
by: Yang, Jian, et al.
Published: (2026)
Researchers of a Feather Flock Together: Endogamy Recruitment Tracks for Industrial Researchers
by: Céline Antonin, et al.
Published: (2025)
by: Céline Antonin, et al.
Published: (2025)
AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size
by: Ostroukhov, Petr, et al.
Published: (2024)
by: Ostroukhov, Petr, et al.
Published: (2024)
Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving
by: Gao, Shihong, et al.
Published: (2025)
by: Gao, Shihong, et al.
Published: (2025)
Partially Observable Gaussian Process Network and Doubly Stochastic Variational Inference
by: Kiroriwal, Saksham, et al.
Published: (2025)
by: Kiroriwal, Saksham, et al.
Published: (2025)
MarginGate: Sparse Margin-Triggered Verification for Batch-Invariant LLM Inference
by: Chu, Kexin, et al.
Published: (2026)
by: Chu, Kexin, et al.
Published: (2026)
Convergence Bound and Critical Batch Size of Muon Optimizer
by: Sato, Naoki, et al.
Published: (2025)
by: Sato, Naoki, et al.
Published: (2025)
No Request Left Behind: Tackling Heterogeneity in Long-Context LLM Inference with Medha
by: Agrawal, Amey, et al.
Published: (2024)
by: Agrawal, Amey, et al.
Published: (2024)
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency
by: Li, Ruixiao, et al.
Published: (2025)
by: Li, Ruixiao, et al.
Published: (2025)
Birds of a Feather Flock Together: Background-Invariant Representations via Linear Structure in VLMs
by: Zaazou, Youssef, et al.
Published: (2026)
by: Zaazou, Youssef, et al.
Published: (2026)
Critical Batch Size Revisited: A Simple Empirical Approach to Large-Batch Language Model Training
by: Merrill, William, et al.
Published: (2025)
by: Merrill, William, et al.
Published: (2025)
Feather: An Elegant Solution to Effective DNN Sparsification
by: Georgoulakis, Athanasios Glentis, et al.
Published: (2023)
by: Georgoulakis, Athanasios Glentis, et al.
Published: (2023)
Inference for Batched Adaptive Experiments
by: Kemper, Jan, et al.
Published: (2025)
by: Kemper, Jan, et al.
Published: (2025)
Sparse Prefix Caching for Hybrid and Recurrent LLM Serving
by: Shirokikh, Mikhail, et al.
Published: (2026)
by: Shirokikh, Mikhail, et al.
Published: (2026)
PrefixWall: Mitigating Prefix Caching Side Channels in Shared LLM Systems
by: Pennas, Panagiotis Georgios, et al.
Published: (2026)
by: Pennas, Panagiotis Georgios, et al.
Published: (2026)
Polar Sparsity: High Throughput Batched LLM Inferencing with Scalable Contextual Sparsity
by: Shrestha, Susav, et al.
Published: (2025)
by: Shrestha, Susav, et al.
Published: (2025)
Surge Phenomenon in Optimal Learning Rate and Batch Size Scaling
by: Li, Shuaipeng, et al.
Published: (2024)
by: Li, Shuaipeng, et al.
Published: (2024)
Smaller Batches, Bigger Gains? Investigating the Impact of Batch Sizes on Reinforcement Learning Based Real-World Production Scheduling
by: Müller, Arthur, et al.
Published: (2024)
by: Müller, Arthur, et al.
Published: (2024)
MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval
by: Srivastava, Saksham Sahai, et al.
Published: (2025)
by: Srivastava, Saksham Sahai, et al.
Published: (2025)
Power Lines: Scaling Laws for Weight Decay and Batch Size in LLM Pre-training
by: Bergsma, Shane, et al.
Published: (2025)
by: Bergsma, Shane, et al.
Published: (2025)
Multi-Bin Batching for Increasing LLM Inference Throughput
by: Guldogan, Ozgur, et al.
Published: (2024)
by: Guldogan, Ozgur, et al.
Published: (2024)
Increasing Batch Size Improves Convergence of Stochastic Gradient Descent with Momentum
by: Kamo, Keisuke, et al.
Published: (2025)
by: Kamo, Keisuke, et al.
Published: (2025)
On the Role of Batch Size in Stochastic Conditional Gradient Methods
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
AntBatchInfer: Elastic Batch Inference in the Kubernetes Cluster
by: Li, Siyuan, et al.
Published: (2024)
by: Li, Siyuan, et al.
Published: (2024)
Birds of a Different Feather Flock Together: Exploring Opportunities and Challenges in Animal-Human-Machine Teaming
by: Cohen, Myke C., et al.
Published: (2025)
by: Cohen, Myke C., et al.
Published: (2025)
PackInfer: Compute- and I/O-Efficient Attention for Batched LLM Inference
by: Ning, Rui, et al.
Published: (2026)
by: Ning, Rui, et al.
Published: (2026)
Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference
by: Recasens, Pol G., et al.
Published: (2025)
by: Recasens, Pol G., et al.
Published: (2025)
AdvPrefix: An Objective for Nuanced LLM Jailbreaks
by: Zhu, Sicheng, et al.
Published: (2024)
by: Zhu, Sicheng, et al.
Published: (2024)
Prefix Grouper: Efficient GRPO Training through Shared-Prefix Forward
by: Liu, Zikang, et al.
Published: (2025)
by: Liu, Zikang, et al.
Published: (2025)
Collaborative Batch Size Optimization for Federated Learning
by: Geimer, Arno, et al.
Published: (2025)
by: Geimer, Arno, et al.
Published: (2025)
Beware of the Batch Size: Hyperparameter Bias in Evaluating LoRA
by: Lee, Sangyoon, et al.
Published: (2026)
by: Lee, Sangyoon, et al.
Published: (2026)
Scaling Law for Language Models Training Considering Batch Size
by: Shuai, Xian, et al.
Published: (2024)
by: Shuai, Xian, et al.
Published: (2024)
Investigating Batch Inference in a Sequential Monte Carlo Framework for Neural Networks
by: Millard, Andrew, et al.
Published: (2026)
by: Millard, Andrew, et al.
Published: (2026)
PrefixQuant: Eliminating Outliers by Prefixed Tokens for Large Language Models Quantization
by: Chen, Mengzhao, et al.
Published: (2024)
by: Chen, Mengzhao, et al.
Published: (2024)
PrefixGPT: Prefix Adder Optimization by a Generative Pre-trained Transformer
by: Ding, Ruogu, et al.
Published: (2025)
by: Ding, Ruogu, et al.
Published: (2025)
Similar Items
-
Reliability Estimation of News Media Sources: Birds of a Feather Flock Together
by: Burdisso, Sergio, et al.
Published: (2024) -
BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching
by: Zheng, Zhen, et al.
Published: (2024) -
Hydragen: High-Throughput LLM Inference with Shared Prefixes
by: Juravsky, Jordan, et al.
Published: (2024) -
Full-Graph vs. Mini-Batch Training: Comprehensive Analysis from a Batch Size and Fan-Out Size Perspective
by: Liu, Mengfan, et al.
Published: (2026) -
SCORPIO: Serving the Right Requests at the Right Time for Heterogeneous SLOs in LLM Inference
by: Tang, Yinghao, et al.
Published: (2025)