Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Metel, Michael R., Chen, Boxing, Rezagholizadeh, Mehdi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Draft on the Fly: Adaptive Self-Speculative Decoding using Cosine Similarity
von: Metel, Michael R., et al.
Veröffentlicht: (2024)
von: Metel, Michael R., et al.
Veröffentlicht: (2024)
On the importance of Data Scale in Pretraining Arabic Language Models
von: Ghaddar, Abbas, et al.
Veröffentlicht: (2024)
von: Ghaddar, Abbas, et al.
Veröffentlicht: (2024)
BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching
von: Zheng, Zhen, et al.
Veröffentlicht: (2024)
von: Zheng, Zhen, et al.
Veröffentlicht: (2024)
Thinking Long, but Short: Stable Sequential Test-Time Scaling for Large Reasoning Models
von: Metel, Michael R., et al.
Veröffentlicht: (2026)
von: Metel, Michael R., et al.
Veröffentlicht: (2026)
X-EcoMLA: Upcycling Pre-Trained Attention into MLA for Efficient and Extreme KV Compression
von: Li, Guihong, et al.
Veröffentlicht: (2025)
von: Li, Guihong, et al.
Veröffentlicht: (2025)
ReGLA: Refining Gated Linear Attention
von: Lu, Peng, et al.
Veröffentlicht: (2025)
von: Lu, Peng, et al.
Veröffentlicht: (2025)
ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
Multi-Bin Batching for Increasing LLM Inference Throughput
von: Guldogan, Ozgur, et al.
Veröffentlicht: (2024)
von: Guldogan, Ozgur, et al.
Veröffentlicht: (2024)
EvolKV: Evolutionary KV Cache Compression for LLM Inference
von: Yu, Bohan, et al.
Veröffentlicht: (2025)
von: Yu, Bohan, et al.
Veröffentlicht: (2025)
SABlock: Semantic-Aware KV Cache Eviction with Adaptive Compression Block Size
von: Chen, Jinhan, et al.
Veröffentlicht: (2025)
von: Chen, Jinhan, et al.
Veröffentlicht: (2025)
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
OTTAWA: Optimal TransporT Adaptive Word Aligner for Hallucination and Omission Translation Errors Detection
von: Huang, Chenyang, et al.
Veröffentlicht: (2024)
von: Huang, Chenyang, et al.
Veröffentlicht: (2024)
BatchGEMBA: Token-Efficient Machine Translation Evaluation with Batched Prompting and Prompt Compression
von: Larionov, Daniil, et al.
Veröffentlicht: (2025)
von: Larionov, Daniil, et al.
Veröffentlicht: (2025)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
CHARP: Conversation History AwaReness Probing for Knowledge-grounded Dialogue Systems
von: Ghaddar, Abbas, et al.
Veröffentlicht: (2024)
von: Ghaddar, Abbas, et al.
Veröffentlicht: (2024)
Sorted LLaMA: Unlocking the Potential of Intermediate Layers of Large Language Models for Dynamic Inference
von: Kavehzadeh, Parsa, et al.
Veröffentlicht: (2023)
von: Kavehzadeh, Parsa, et al.
Veröffentlicht: (2023)
KV-Compress: Paged KV-Cache Compression with Variable Compression Rates per Attention Head
von: Rehg, Isaac
Veröffentlicht: (2024)
von: Rehg, Isaac
Veröffentlicht: (2024)
Are Optimal Algorithms Still Optimal? Rethinking Sorting in LLM-Based Pairwise Ranking with Batching and Caching
von: Wisznia, Juan, et al.
Veröffentlicht: (2025)
von: Wisznia, Juan, et al.
Veröffentlicht: (2025)
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity
von: Ma, Da, et al.
Veröffentlicht: (2024)
von: Ma, Da, et al.
Veröffentlicht: (2024)
NestedKV: Nested Memory Routing for Long-Context KV Cache Compression
von: Chen, Hong, et al.
Veröffentlicht: (2026)
von: Chen, Hong, et al.
Veröffentlicht: (2026)
CaliDrop: KV Cache Compression with Calibration
von: Su, Yi, et al.
Veröffentlicht: (2025)
von: Su, Yi, et al.
Veröffentlicht: (2025)
PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference
von: Yang, Dongjie, et al.
Veröffentlicht: (2024)
von: Yang, Dongjie, et al.
Veröffentlicht: (2024)
Lossless KV Cache Compression to 2%
von: Yang, Zhen, et al.
Veröffentlicht: (2024)
von: Yang, Zhen, et al.
Veröffentlicht: (2024)
KV-CoRE: Benchmarking Data-Dependent Low-Rank Compressibility of KV-Caches in LLMs
von: Chen, Jian, et al.
Veröffentlicht: (2026)
von: Chen, Jian, et al.
Veröffentlicht: (2026)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
CodeComp: Structural KV Cache Compression for Agentic Coding
von: Chen, Qiujiang, et al.
Veröffentlicht: (2026)
von: Chen, Qiujiang, et al.
Veröffentlicht: (2026)
EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction
von: Ji, Shiyu, et al.
Veröffentlicht: (2026)
von: Ji, Shiyu, et al.
Veröffentlicht: (2026)
RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression
von: Behnam, Payman, et al.
Veröffentlicht: (2025)
von: Behnam, Payman, et al.
Veröffentlicht: (2025)
SemantiCache: Efficient KV Cache Compression via Semantic Chunking and Clustered Merging
von: Wu, Shunlong, et al.
Veröffentlicht: (2026)
von: Wu, Shunlong, et al.
Veröffentlicht: (2026)
KVSculpt: KV Cache Compression as Distillation
von: Jiang, Bo, et al.
Veröffentlicht: (2026)
von: Jiang, Bo, et al.
Veröffentlicht: (2026)
R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs
von: Zhou, Xiabin, et al.
Veröffentlicht: (2024)
von: Zhou, Xiabin, et al.
Veröffentlicht: (2024)
One Size Does Not Fit All: Token-Wise Adaptive Compression for KV Cache
von: Lu, Liming, et al.
Veröffentlicht: (2026)
von: Lu, Liming, et al.
Veröffentlicht: (2026)
PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference
von: Patel, Ishan, et al.
Veröffentlicht: (2026)
von: Patel, Ishan, et al.
Veröffentlicht: (2026)
AttentionPredictor: Temporal Patterns Matter for KV Cache Compression
von: Yang, Qingyue, et al.
Veröffentlicht: (2025)
von: Yang, Qingyue, et al.
Veröffentlicht: (2025)
SmolTulu: Higher Learning Rate to Batch Size Ratios Can Lead to Better Reasoning in SLMs
von: Alrashed, Sultan
Veröffentlicht: (2024)
von: Alrashed, Sultan
Veröffentlicht: (2024)
S2D: Sorted Speculative Decoding For More Efficient Deployment of Nested Large Language Models
von: Kavehzadeh, Parsa, et al.
Veröffentlicht: (2024)
von: Kavehzadeh, Parsa, et al.
Veröffentlicht: (2024)
Key, Value, Compress: A Systematic Exploration of KV Cache Compression Techniques
von: Javidnia, Neusha, et al.
Veröffentlicht: (2025)
von: Javidnia, Neusha, et al.
Veröffentlicht: (2025)
HeteroCache: A Dynamic Retrieval Approach to Heterogeneous KV Cache Compression for Long-Context LLM Inference
von: Shi, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Shi, Zhiyuan, et al.
Veröffentlicht: (2026)
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
von: Cai, Zefan, et al.
Veröffentlicht: (2024)
von: Cai, Zefan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Draft on the Fly: Adaptive Self-Speculative Decoding using Cosine Similarity
von: Metel, Michael R., et al.
Veröffentlicht: (2024) -
On the importance of Data Scale in Pretraining Arabic Language Models
von: Ghaddar, Abbas, et al.
Veröffentlicht: (2024) -
BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching
von: Zheng, Zhen, et al.
Veröffentlicht: (2024) -
Thinking Long, but Short: Stable Sequential Test-Time Scaling for Large Reasoning Models
von: Metel, Michael R., et al.
Veröffentlicht: (2026) -
X-EcoMLA: Upcycling Pre-Trained Attention into MLA for Efficient and Extreme KV Compression
von: Li, Guihong, et al.
Veröffentlicht: (2025)