VeriCache: Turning Lossy KV Cache into Lossless LLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Yao, Jiayi, Shen, Samuel, Du, Kuntai, Feng, Shaoting, Seo, Dongjoo, Zhang, Rui, Huang, Yuyang, Liu, Yuhan, Lu, Shan, Jiang, Junchen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Comparative Characterization of KV Cache Management Strategies for LLM Inference
by: Mamo, Oteo, et al.
Published: (2026)
by: Mamo, Oteo, et al.
Published: (2026)
Titanus: Enabling KV Cache Pruning and Quantization On-the-Fly for LLM Acceleration
by: Chen, Peilin, et al.
Published: (2025)
by: Chen, Peilin, et al.
Published: (2025)
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
by: Fang, Yunhua, et al.
Published: (2025)
by: Fang, Yunhua, et al.
Published: (2025)
UniCAIM: A Unified CAM/CIM Architecture with Static-Dynamic KV Cache Pruning for Efficient Long-Context LLM Inference
by: Xu, Weikai, et al.
Published: (2025)
by: Xu, Weikai, et al.
Published: (2025)
Towards More Economical Context-Augmented LLM Generation by Reusing Stored KV Cache
by: Li, Hanchen, et al.
Published: (2025)
by: Li, Hanchen, et al.
Published: (2025)
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator
by: Wang, Zhican, et al.
Published: (2025)
by: Wang, Zhican, et al.
Published: (2025)
Adaptive Multi-Objective Tiered Storage Configuration for KV Cache in LLM Service
by: Zheng, Xianzhe, et al.
Published: (2026)
by: Zheng, Xianzhe, et al.
Published: (2026)
Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
by: Kim, Minsu, et al.
Published: (2025)
by: Kim, Minsu, et al.
Published: (2025)
A Joint Learning Approach to Hardware Caching and Prefetching
by: Yuan, Samuel, et al.
Published: (2025)
by: Yuan, Samuel, et al.
Published: (2025)
Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge Computing
by: Xia, Tianhua, et al.
Published: (2025)
by: Xia, Tianhua, et al.
Published: (2025)
Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache Management Beyond GPU Limits
by: Kim, Dowon, et al.
Published: (2025)
by: Kim, Dowon, et al.
Published: (2025)
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
by: li, Fei, et al.
Published: (2026)
by: li, Fei, et al.
Published: (2026)
LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference
by: Liu, Yuhan, et al.
Published: (2025)
by: Liu, Yuhan, et al.
Published: (2025)
CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving
by: Liu, Yuhan, et al.
Published: (2023)
by: Liu, Yuhan, et al.
Published: (2023)
PREFENDER: A Prefetching Defender against Cache Side Channel Attacks as A Pretender
by: Li, Luyi, et al.
Published: (2023)
by: Li, Luyi, et al.
Published: (2023)
Bit-Flip Vulnerability of Shared KV-Cache Blocks in LLM Serving Systems
by: Yamamoto, Yuji, et al.
Published: (2026)
by: Yamamoto, Yuji, et al.
Published: (2026)
Fletch: File-System Metadata Caching in Programmable Switches
by: Liu, Qingxiu, et al.
Published: (2025)
by: Liu, Qingxiu, et al.
Published: (2025)
Cache Your Prompt When It's Green: Carbon-Aware Caching for Large Language Model Serving
by: Tian, Yuyang, et al.
Published: (2025)
by: Tian, Yuyang, et al.
Published: (2025)
VeriRAG: A Retrieval-Augmented Framework for Automated RTL Testability Repair
by: Qi, Haomin, et al.
Published: (2025)
by: Qi, Haomin, et al.
Published: (2025)
PiKV: KV Cache Management System for Mixture of Experts
by: Liu, Dong, et al.
Published: (2025)
by: Liu, Dong, et al.
Published: (2025)
BackCache: Mitigating Contention-Based Cache Timing Attacks by Hiding Cache Line Evictions
by: Wang, Quancheng, et al.
Published: (2023)
by: Wang, Quancheng, et al.
Published: (2023)
PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
by: Yüzügüler, Ahmet Caner, et al.
Published: (2025)
by: Yüzügüler, Ahmet Caner, et al.
Published: (2025)
Exploring DRAM Cache Prefetching for Pooled Memory
by: Tirumalasetty, Chandrahas, et al.
Published: (2024)
by: Tirumalasetty, Chandrahas, et al.
Published: (2024)
Potential and Limitation of High-Frequency Cores and Caches
by: Pai, Kunal, et al.
Published: (2024)
by: Pai, Kunal, et al.
Published: (2024)
TDRAM: Tag-enhanced DRAM for Efficient Caching
by: Babaie, Maryam, et al.
Published: (2024)
by: Babaie, Maryam, et al.
Published: (2024)
Adaptive Cache Pollution Control for Large Language Model Inference Workloads Using Temporal CNN-Based Prediction and Priority-Aware Replacement
by: Liu, Songze, et al.
Published: (2025)
by: Liu, Songze, et al.
Published: (2025)
The Avatar Cache: Enabling On-Demand Security with Morphable Cache Architecture
by: Bhatla, Anubhav, et al.
Published: (2026)
by: Bhatla, Anubhav, et al.
Published: (2026)
Random and Safe Cache Architecture to Defeat Cache Timing Attacks
by: Hu, Guangyuan, et al.
Published: (2023)
by: Hu, Guangyuan, et al.
Published: (2023)
Systematic Evaluation of Randomized Cache Designs against Cache Occupancy
by: Chakraborty, Anirban, et al.
Published: (2023)
by: Chakraborty, Anirban, et al.
Published: (2023)
DCI: A Coordinated Allocation and Filling Workload-Aware Dual-Cache Allocation GNN Inference Acceleration System
by: Luo, Yi, et al.
Published: (2025)
by: Luo, Yi, et al.
Published: (2025)
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference
by: Yubeaton, Patrick, et al.
Published: (2025)
by: Yubeaton, Patrick, et al.
Published: (2025)
Improving the Representativeness of Simulation Intervals for the Cache Memory System
by: Bueno, Nicolas, et al.
Published: (2024)
by: Bueno, Nicolas, et al.
Published: (2024)
SliceMoE: Bit-Sliced Expert Caching under Miss-Rate Constraints for Efficient MoE Inference
by: Choi, Yuseon, et al.
Published: (2025)
by: Choi, Yuseon, et al.
Published: (2025)
MERE: Hardware-Software Co-Design for Masking Cache Miss Latency in Embedded Processors
by: You, Dean, et al.
Published: (2025)
by: You, Dean, et al.
Published: (2025)
Multi-Dimensional Vector ISA Extension for Mobile In-Cache Computing
by: Khadem, Alireza, et al.
Published: (2025)
by: Khadem, Alireza, et al.
Published: (2025)
Pickle Prefetcher: Programmable and Scalable Last-Level Cache Prefetcher
by: Nguyen, Hoa, et al.
Published: (2025)
by: Nguyen, Hoa, et al.
Published: (2025)
CMD: A Cache-assisted GPU Memory Deduplication Architecture
by: Zhao, Wei, et al.
Published: (2024)
by: Zhao, Wei, et al.
Published: (2024)
Bandwidth-Effective DRAM Cache for GPUs with Storage-Class Memory
by: Hong, Jeongmin, et al.
Published: (2024)
by: Hong, Jeongmin, et al.
Published: (2024)
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
by: Du, Dayou, et al.
Published: (2025)
by: Du, Dayou, et al.
Published: (2025)
RollingCache: Using Runtime Behavior to Defend Against Cache Side Channel Attacks
by: Ojha, Divya, et al.
Published: (2024)
by: Ojha, Divya, et al.
Published: (2024)
Similar Items
-
Comparative Characterization of KV Cache Management Strategies for LLM Inference
by: Mamo, Oteo, et al.
Published: (2026) -
Titanus: Enabling KV Cache Pruning and Quantization On-the-Fly for LLM Acceleration
by: Chen, Peilin, et al.
Published: (2025) -
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
by: Fang, Yunhua, et al.
Published: (2025) -
UniCAIM: A Unified CAM/CIM Architecture with Static-Dynamic KV Cache Pruning for Efficient Long-Context LLM Inference
by: Xu, Weikai, et al.
Published: (2025) -
Towards More Economical Context-Augmented LLM Generation by Reusing Stored KV Cache
by: Li, Hanchen, et al.
Published: (2025)