CachePrune: Teaching LLMs What Not to Follow via KV-Cache Editing
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Rui, Wu, Junda, Xia, Yu, Yu, Tong, Zhang, Ruiyi, Rossi, Ryan, Mitra, Subrata, Yao, Lina, McAuley, Julian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CachePrune: Privacy-Aware and Fine-Grained KV Cache Sharing for Efficient LLM Inference
by: Wu, Guanlong, et al.
Published: (2026)
by: Wu, Guanlong, et al.
Published: (2026)
Pluralistic Off-policy Evaluation and Alignment
by: Huang, Chengkai, et al.
Published: (2025)
by: Huang, Chengkai, et al.
Published: (2025)
DICE: Dynamic In-Context Example Selection in LLM Agents via Efficient Knowledge Transfer
by: Wang, Ruoyu, et al.
Published: (2025)
by: Wang, Ruoyu, et al.
Published: (2025)
CacheProbe: Auditing Prompt Cache Isolation in Gateway APIs
by: Fahey, Ryan
Published: (2026)
by: Fahey, Ryan
Published: (2026)
Cachemir: Fully Homomorphic Encrypted Inference of Generative Large Language Model with KV Cache
by: Yu, Ye, et al.
Published: (2026)
by: Yu, Ye, et al.
Published: (2026)
CryptoGen: Secure Transformer Generation with Encrypted KV-Cache Reuse
by: Zhang, Hedong, et al.
Published: (2026)
by: Zhang, Hedong, et al.
Published: (2026)
Whose Narrative is it Anyway? A KV Cache Manipulation Attack
by: Ganesh, Mukkesh, et al.
Published: (2025)
by: Ganesh, Mukkesh, et al.
Published: (2025)
MPCache: MPC-Friendly KV Cache Eviction for Efficient Private LLM Inference
by: Zeng, Wenxuan, et al.
Published: (2025)
by: Zeng, Wenxuan, et al.
Published: (2025)
SAND: Boosting LLM Agents with Self-Taught Action Deliberation
by: Xia, Yu, et al.
Published: (2025)
by: Xia, Yu, et al.
Published: (2025)
Weakly-supervised VLM-guided Partial Contrastive Learning for Visual Language Navigation
by: Wang, Ruoyu, et al.
Published: (2025)
by: Wang, Ruoyu, et al.
Published: (2025)
Visual Prompting in Multimodal Large Language Models: A Survey
by: Wu, Junda, et al.
Published: (2024)
by: Wu, Junda, et al.
Published: (2024)
Bit-Flip Vulnerability of Shared KV-Cache Blocks in LLM Serving Systems
by: Yamamoto, Yuji, et al.
Published: (2026)
by: Yamamoto, Yuji, et al.
Published: (2026)
Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference
by: Chu, Kexin, et al.
Published: (2025)
by: Chu, Kexin, et al.
Published: (2025)
Shadow in the Cache: Unveiling and Mitigating Privacy Risks of KV-cache in LLM Inference
by: Luo, Zhifan, et al.
Published: (2025)
by: Luo, Zhifan, et al.
Published: (2025)
Addendum: Systematic Evaluation of Randomized Cache Designs against Cache Occupancy
by: Chakraborty, Anirban, et al.
Published: (2025)
by: Chakraborty, Anirban, et al.
Published: (2025)
BackCache: Mitigating Contention-Based Cache Timing Attacks by Hiding Cache Line Evictions
by: Wang, Quancheng, et al.
Published: (2023)
by: Wang, Quancheng, et al.
Published: (2023)
Knowledge-Aware Query Expansion with Large Language Models for Textual and Relational Retrieval
by: Xia, Yu, et al.
Published: (2024)
by: Xia, Yu, et al.
Published: (2024)
Listwise Preference Diffusion Optimization for User Behavior Trajectories Prediction
by: Huang, Hongtao, et al.
Published: (2025)
by: Huang, Hongtao, et al.
Published: (2025)
Towards Agentic Recommender Systems in the Era of Multimodal Large Language Models
by: Huang, Chengkai, et al.
Published: (2025)
by: Huang, Chengkai, et al.
Published: (2025)
The Avatar Cache: Enabling On-Demand Security with Morphable Cache Architecture
by: Bhatla, Anubhav, et al.
Published: (2026)
by: Bhatla, Anubhav, et al.
Published: (2026)
Random and Safe Cache Architecture to Defeat Cache Timing Attacks
by: Hu, Guangyuan, et al.
Published: (2023)
by: Hu, Guangyuan, et al.
Published: (2023)
Systematic Evaluation of Randomized Cache Designs against Cache Occupancy
by: Chakraborty, Anirban, et al.
Published: (2023)
by: Chakraborty, Anirban, et al.
Published: (2023)
Towards Threshold-Free KV Cache Pruning
by: Ni, Xuanfan, et al.
Published: (2025)
by: Ni, Xuanfan, et al.
Published: (2025)
Hidden Web Caches Discovery
by: Golinelli, Matteo, et al.
Published: (2024)
by: Golinelli, Matteo, et al.
Published: (2024)
Multi-Agent Collaborative Filtering: Orchestrating Users and Items for Agentic Recommendations
by: Xia, Yu, et al.
Published: (2025)
by: Xia, Yu, et al.
Published: (2025)
CacheTrap: Unveiling a Stealthier Gray-Box Trojan against LLMs
by: Nahian, Mohaiminul Al, et al.
Published: (2025)
by: Nahian, Mohaiminul Al, et al.
Published: (2025)
RedVisor: Reasoning-Aware Prompt Injection Defense via Zero-Copy KV Cache Reuse
by: Liu, Mingrui, et al.
Published: (2026)
by: Liu, Mingrui, et al.
Published: (2026)
Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs
by: Bui, Ngoc, et al.
Published: (2025)
by: Bui, Ngoc, et al.
Published: (2025)
RollingCache: Using Runtime Behavior to Defend Against Cache Side Channel Attacks
by: Ojha, Divya, et al.
Published: (2024)
by: Ojha, Divya, et al.
Published: (2024)
Active Learning for Direct Preference Optimization
by: Kveton, Branislav, et al.
Published: (2025)
by: Kveton, Branislav, et al.
Published: (2025)
PCG: Mitigating Conflict-based Cache Side-channel Attacks with Prefetching
by: Jiang, Fang, et al.
Published: (2024)
by: Jiang, Fang, et al.
Published: (2024)
Attacks on Approximate Caches in Text-to-Image Diffusion Models
by: Sun, Desen, et al.
Published: (2025)
by: Sun, Desen, et al.
Published: (2025)
Embedding-Informed Adaptive Retrieval-Augmented Generation of Large Language Models
by: Huang, Chengkai, et al.
Published: (2024)
by: Huang, Chengkai, et al.
Published: (2024)
Federated Large Language Models: Current Progress and Future Directions
by: Yao, Yuhang, et al.
Published: (2024)
by: Yao, Yuhang, et al.
Published: (2024)
EXAM: Exploiting Exclusive System-Level Cache in Apple M-Series SoCs for Enhanced Cache Occupancy Attacks
by: Xu, Tianhong, et al.
Published: (2025)
by: Xu, Tianhong, et al.
Published: (2025)
Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation
by: Agarwal, Shubham, et al.
Published: (2025)
by: Agarwal, Shubham, et al.
Published: (2025)
CacheSquash: Making caches speculation-aware
by: ElAtali, Hossam, et al.
Published: (2024)
by: ElAtali, Hossam, et al.
Published: (2024)
Prime+Retouch: When Cache is Locked and Leaked
by: Lee, Jaehyuk, et al.
Published: (2024)
by: Lee, Jaehyuk, et al.
Published: (2024)
Systematic Assessment of Cache Timing Vulnerabilities on RISC-V Processors
by: Austa, Cédrick, et al.
Published: (2025)
by: Austa, Cédrick, et al.
Published: (2025)
dKV-Cache: The Cache for Diffusion Language Models
by: Ma, Xinyin, et al.
Published: (2025)
by: Ma, Xinyin, et al.
Published: (2025)
Similar Items
-
CachePrune: Privacy-Aware and Fine-Grained KV Cache Sharing for Efficient LLM Inference
by: Wu, Guanlong, et al.
Published: (2026) -
Pluralistic Off-policy Evaluation and Alignment
by: Huang, Chengkai, et al.
Published: (2025) -
DICE: Dynamic In-Context Example Selection in LLM Agents via Efficient Knowledge Transfer
by: Wang, Ruoyu, et al.
Published: (2025) -
CacheProbe: Auditing Prompt Cache Isolation in Gateway APIs
by: Fahey, Ryan
Published: (2026) -
Cachemir: Fully Homomorphic Encrypted Inference of Generative Large Language Model with KV Cache
by: Yu, Ye, et al.
Published: (2026)