Can Increasing the Hit Ratio Hurt Cache Throughput? (Long Version)
Fuente:
arXiv
Saved in:
| Main Authors: | Qiu, Ziyue, Yang, Juncheng, Harchol-Balter, Mor |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Upper Bound on the M/M/k Queue With Deterministic Setup Times
by: Williams, Jalani, et al.
Published: (2025)
by: Williams, Jalani, et al.
Published: (2025)
Analysis of Markovian Arrivals and Service with Applications to Intermittent Overload
by: Grosof, Isaac, et al.
Published: (2024)
by: Grosof, Isaac, et al.
Published: (2024)
SPLIT: SymPathy for Large jobs Improves Tail latency
by: Li, Zhouzi, et al.
Published: (2026)
by: Li, Zhouzi, et al.
Published: (2026)
Improving Upon the generalized c-mu rule: a Whittle approach
by: Li, Zhouzi, et al.
Published: (2025)
by: Li, Zhouzi, et al.
Published: (2025)
LookAhead: The Optimal Non-decreasing Index Policy for a Time-Varying Holding Cost problem
by: Gurushankar, Keerthana, et al.
Published: (2026)
by: Gurushankar, Keerthana, et al.
Published: (2026)
Asymptotically Optimal Scheduling of Multiple Parallelizable Job Classes
by: Berg, Benjamin, et al.
Published: (2024)
by: Berg, Benjamin, et al.
Published: (2024)
How to Rent GPUs on a Budget
by: Li, Zhouzi, et al.
Published: (2024)
by: Li, Zhouzi, et al.
Published: (2024)
Mean field optimal Core Allocation across Malleable jobs
by: Li, Zhouzi, et al.
Published: (2026)
by: Li, Zhouzi, et al.
Published: (2026)
A Zoned Storage Optimized Flash Cache on ZNS SSDs
by: Yang, Chongzhuo, et al.
Published: (2024)
by: Yang, Chongzhuo, et al.
Published: (2024)
Introducing the Arm-membench Throughput Benchmark
by: Burth, Cyrill, et al.
Published: (2025)
by: Burth, Cyrill, et al.
Published: (2025)
End-to-End Throughput Benchmarking of Portable Deterministic CNN-Based Signal Processing Pipelines
by: Boerkamp, Christiaan, et al.
Published: (2026)
by: Boerkamp, Christiaan, et al.
Published: (2026)
XRFlux: Virtual Reality Benchmark for Edge Caching Systems
by: Alfares, Nader, et al.
Published: (2024)
by: Alfares, Nader, et al.
Published: (2024)
SparseX: Efficient Segment-Level KV Cache Sharing for Interleaved LLM Serving
by: Zhang, Quqing, et al.
Published: (2026)
by: Zhang, Quqing, et al.
Published: (2026)
One-Hop Sub-Query Result Caches for Graph Database Systems
by: Nguyen, Hieu, et al.
Published: (2024)
by: Nguyen, Hieu, et al.
Published: (2024)
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
by: Du, Dayou, et al.
Published: (2025)
by: Du, Dayou, et al.
Published: (2025)
The Bicameral Cache: a split cache for vector architectures
by: Rebolledo, Susana, et al.
Published: (2024)
by: Rebolledo, Susana, et al.
Published: (2024)
2DIO: A Cache-Accurate Storage Microbenchmark
by: Wang, Yirong, et al.
Published: (2026)
by: Wang, Yirong, et al.
Published: (2026)
Enhancing Instruction Prefetching via Cache and TLB Management
by: Jamet, Alexandre Valentin, et al.
Published: (2026)
by: Jamet, Alexandre Valentin, et al.
Published: (2026)
Dual-Select FMA Butterfly for FFT: Eliminating Twiddle Factor Singularities with Bounded Precomputed Ratios
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
Long-term Monitoring of Kernel and Hardware Events to Understand Latency Variance
by: Zhou, Fang, et al.
Published: (2026)
by: Zhou, Fang, et al.
Published: (2026)
Analysis and Evaluation of Using Microsecond-Latency Memory for In-Memory Indices and Caches in SSD-Based Key-Value Stores
by: Bando, Yosuke, et al.
Published: (2025)
by: Bando, Yosuke, et al.
Published: (2025)
Leveraging Approximate Caching for Faster Retrieval-Augmented Generation
by: Bergman, Shai, et al.
Published: (2025)
by: Bergman, Shai, et al.
Published: (2025)
Faster LLM Inference using DBMS-Inspired Preemption and Cache Replacement Policies
by: Kim, Kyoungmin, et al.
Published: (2024)
by: Kim, Kyoungmin, et al.
Published: (2024)
GPU-Accelerated INT8 Quantization for KV Cache Compression in Large Language Models
by: Taneja, Maanas, et al.
Published: (2026)
by: Taneja, Maanas, et al.
Published: (2026)
BOA Constrictor: Squeezing Performance out of GPUs in the Cloud via Budget-Optimal Allocation
by: Li, Zhouzi, et al.
Published: (2026)
by: Li, Zhouzi, et al.
Published: (2026)
Optimizing CPU Cache Utilization in Cloud VMs with Accurate Cache Abstraction
by: Tofigh, Mani, et al.
Published: (2025)
by: Tofigh, Mani, et al.
Published: (2025)
When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
DistZO2: High-Throughput and Memory-Efficient Zeroth-Order Fine-tuning LLMs with Distributed Parallel Computing
by: Wang, Liangyu, et al.
Published: (2025)
by: Wang, Liangyu, et al.
Published: (2025)
MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
by: Xue, Leyang, et al.
Published: (2024)
by: Xue, Leyang, et al.
Published: (2024)
How to Increase Energy Efficiency with a Single Linux Command
by: Jelvani, Alborz, et al.
Published: (2025)
by: Jelvani, Alborz, et al.
Published: (2025)
Using Evolutionary Algorithms to Find Cache-Friendly Generalized Morton Layouts for Arrays
by: Swatman, Stephen Nicholas, et al.
Published: (2023)
by: Swatman, Stephen Nicholas, et al.
Published: (2023)
Impact of Data-Oriented and Object-Oriented Design on Performance and Cache Utilization with Artificial Intelligence Algorithms in Multi-Threaded CPUs
by: Arantes, Gabriel M., et al.
Published: (2025)
by: Arantes, Gabriel M., et al.
Published: (2025)
Cache Blocking of Distributed-Memory Parallel Matrix Power Kernels
by: Lacey, Dane C., et al.
Published: (2024)
by: Lacey, Dane C., et al.
Published: (2024)
Hold Onto That Thought: Assessing KV Cache Compression On Reasoning
by: Liu, Minghui, et al.
Published: (2025)
by: Liu, Minghui, et al.
Published: (2025)
OrbitFlow: SLO-Aware Long-Context LLM Serving with Fine-Grained KV Cache Reconfiguration
by: Ma, Xinyue, et al.
Published: (2026)
by: Ma, Xinyue, et al.
Published: (2026)
HiKonv: Maximizing the Throughput of Quantized Convolution With Novel Bit-wise Management and Computation
by: Chen, Yao, et al.
Published: (2022)
by: Chen, Yao, et al.
Published: (2022)
Adaptive Cache Pollution Control for Large Language Model Inference Workloads Using Temporal CNN-Based Prediction and Priority-Aware Replacement
by: Liu, Songze, et al.
Published: (2025)
by: Liu, Songze, et al.
Published: (2025)
H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference
by: Fu, Zizhuo, et al.
Published: (2025)
by: Fu, Zizhuo, et al.
Published: (2025)
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
by: Liu, Zirui, et al.
Published: (2024)
by: Liu, Zirui, et al.
Published: (2024)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
by: Liu, Guangda, et al.
Published: (2024)
by: Liu, Guangda, et al.
Published: (2024)
Similar Items
-
An Upper Bound on the M/M/k Queue With Deterministic Setup Times
by: Williams, Jalani, et al.
Published: (2025) -
Analysis of Markovian Arrivals and Service with Applications to Intermittent Overload
by: Grosof, Isaac, et al.
Published: (2024) -
SPLIT: SymPathy for Large jobs Improves Tail latency
by: Li, Zhouzi, et al.
Published: (2026) -
Improving Upon the generalized c-mu rule: a Whittle approach
by: Li, Zhouzi, et al.
Published: (2025) -
LookAhead: The Optimal Non-decreasing Index Policy for a Time-Varying Holding Cost problem
by: Gurushankar, Keerthana, et al.
Published: (2026)