Probing the Prompt KV Cache: Where It Becomes Dispensable
Fuente:
arXiv
Salvato in:
| Autori principali: | Kumar, Vinayshekhar Bannihatti, Arivazhagan, Manoj Ghuhan, Makhija, Disha, Gangadharaiah, Rashmi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Syntax Without Semantics: Teaching Large Language Models to Code in an Unseen Language
di: Kumar, Vinayshekhar Bannihatti, et al.
Pubblicazione: (2026)
di: Kumar, Vinayshekhar Bannihatti, et al.
Pubblicazione: (2026)
Neural Breadcrumbs: Membership Inference Attacks on LLMs Through Hidden State and Attention Pattern Analysis
di: Makhija, Disha, et al.
Pubblicazione: (2025)
di: Makhija, Disha, et al.
Pubblicazione: (2025)
When Facts Change: Probing LLMs on Evolving Knowledge with evolveQA
di: Nakshatri, Nishanth Sridhar, et al.
Pubblicazione: (2025)
di: Nakshatri, Nishanth Sridhar, et al.
Pubblicazione: (2025)
Multi-Faceted Evaluation of Tool-Augmented Dialogue Systems
di: Hou, Zhaoyi Joey, et al.
Pubblicazione: (2025)
di: Hou, Zhaoyi Joey, et al.
Pubblicazione: (2025)
FairGen: Controlling Sensitive Attributes for Fair Generations in Diffusion Models via Adaptive Latent Guidance
di: Kang, Mintong, et al.
Pubblicazione: (2025)
di: Kang, Mintong, et al.
Pubblicazione: (2025)
Steering Without Breaking: Mechanistically Informed Interventions for Discrete Diffusion Language Models
di: Zhou, Hanhan, et al.
Pubblicazione: (2026)
di: Zhou, Hanhan, et al.
Pubblicazione: (2026)
Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks
di: Wang, Zheng, et al.
Pubblicazione: (2024)
di: Wang, Zheng, et al.
Pubblicazione: (2024)
dKV-Cache: The Cache for Diffusion Language Models
di: Ma, Xinyin, et al.
Pubblicazione: (2025)
di: Ma, Xinyin, et al.
Pubblicazione: (2025)
Bring Your Own KG: Self-Supervised Program Synthesis for Zero-Shot KGQA
di: Agarwal, Dhruv, et al.
Pubblicazione: (2023)
di: Agarwal, Dhruv, et al.
Pubblicazione: (2023)
Where Matters More Than What: Decoding-aligned KV Cache Compression via Position-aware Pseudo Queries
di: Tian, Zhenxu, et al.
Pubblicazione: (2026)
di: Tian, Zhenxu, et al.
Pubblicazione: (2026)
NestedKV: Nested Memory Routing for Long-Context KV Cache Compression
di: Chen, Hong, et al.
Pubblicazione: (2026)
di: Chen, Hong, et al.
Pubblicazione: (2026)
EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction
di: Ji, Shiyu, et al.
Pubblicazione: (2026)
di: Ji, Shiyu, et al.
Pubblicazione: (2026)
CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective
di: Feng, Yuan, et al.
Pubblicazione: (2025)
di: Feng, Yuan, et al.
Pubblicazione: (2025)
RefreshKV: Updating Small KV Cache During Long-form Generation
di: Xu, Fangyuan, et al.
Pubblicazione: (2024)
di: Xu, Fangyuan, et al.
Pubblicazione: (2024)
CaliDrop: KV Cache Compression with Calibration
di: Su, Yi, et al.
Pubblicazione: (2025)
di: Su, Yi, et al.
Pubblicazione: (2025)
EvolKV: Evolutionary KV Cache Compression for LLM Inference
di: Yu, Bohan, et al.
Pubblicazione: (2025)
di: Yu, Bohan, et al.
Pubblicazione: (2025)
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs
di: Zhou, Xiabin, et al.
Pubblicazione: (2024)
di: Zhou, Xiabin, et al.
Pubblicazione: (2024)
GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness
di: Huang, Kung-Hsiang, et al.
Pubblicazione: (2025)
di: Huang, Kung-Hsiang, et al.
Pubblicazione: (2025)
GraphKV: Breaking the Static Selection Paradigm with Graph-Based KV Cache Eviction
di: Li, Xuelin, et al.
Pubblicazione: (2025)
di: Li, Xuelin, et al.
Pubblicazione: (2025)
Lossless KV Cache Compression to 2%
di: Yang, Zhen, et al.
Pubblicazione: (2024)
di: Yang, Zhen, et al.
Pubblicazione: (2024)
G-KV: Decoding-Time KV Cache Eviction with Global Attention
di: Liao, Mengqi, et al.
Pubblicazione: (2025)
di: Liao, Mengqi, et al.
Pubblicazione: (2025)
R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
di: Cai, Zefan, et al.
Pubblicazione: (2025)
di: Cai, Zefan, et al.
Pubblicazione: (2025)
SemantiCache: Efficient KV Cache Compression via Semantic Chunking and Clustered Merging
di: Wu, Shunlong, et al.
Pubblicazione: (2026)
di: Wu, Shunlong, et al.
Pubblicazione: (2026)
EntropyCache: Decoded Token Entropy Guided KV Caching for Diffusion Language Models
di: Cheong, Minsoo, et al.
Pubblicazione: (2026)
di: Cheong, Minsoo, et al.
Pubblicazione: (2026)
QAQ: Quality Adaptive Quantization for LLM KV Cache
di: Dong, Shichen, et al.
Pubblicazione: (2024)
di: Dong, Shichen, et al.
Pubblicazione: (2024)
Accurate KV Cache Quantization with Outlier Tokens Tracing
di: Su, Yi, et al.
Pubblicazione: (2025)
di: Su, Yi, et al.
Pubblicazione: (2025)
Taming the Fragility of KV Cache Eviction in LLM Inference
di: Feng, Yuan, et al.
Pubblicazione: (2025)
di: Feng, Yuan, et al.
Pubblicazione: (2025)
Constrained Decoding with Speculative Lookaheads
di: Nakshatri, Nishanth, et al.
Pubblicazione: (2024)
di: Nakshatri, Nishanth, et al.
Pubblicazione: (2024)
DASH-KV: Accelerating Long-Context LLM Inference via Asymmetric KV Cache Hashing
di: Guo, Jinyu, et al.
Pubblicazione: (2026)
di: Guo, Jinyu, et al.
Pubblicazione: (2026)
KV-CoRE: Benchmarking Data-Dependent Low-Rank Compressibility of KV-Caches in LLMs
di: Chen, Jian, et al.
Pubblicazione: (2026)
di: Chen, Jian, et al.
Pubblicazione: (2026)
KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding
di: Shi, Luohe, et al.
Pubblicazione: (2025)
di: Shi, Luohe, et al.
Pubblicazione: (2025)
KV-Compress: Paged KV-Cache Compression with Variable Compression Rates per Attention Head
di: Rehg, Isaac
Pubblicazione: (2024)
di: Rehg, Isaac
Pubblicazione: (2024)
Task-KV: Task-aware KV Cache Optimization via Semantic Differentiation of Attention Heads
di: He, Xingyang, et al.
Pubblicazione: (2025)
di: He, Xingyang, et al.
Pubblicazione: (2025)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
di: Liu, Xiang, et al.
Pubblicazione: (2025)
di: Liu, Xiang, et al.
Pubblicazione: (2025)
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
di: Cai, Zefan, et al.
Pubblicazione: (2024)
di: Cai, Zefan, et al.
Pubblicazione: (2024)
Towards Threshold-Free KV Cache Pruning
di: Ni, Xuanfan, et al.
Pubblicazione: (2025)
di: Ni, Xuanfan, et al.
Pubblicazione: (2025)
KV Cache Steering for Controlling Frozen LLMs
di: Belitsky, Max, et al.
Pubblicazione: (2025)
di: Belitsky, Max, et al.
Pubblicazione: (2025)
Where Do Self-Supervised Speech Models Become Unfair?
di: Herron, Felix, et al.
Pubblicazione: (2026)
di: Herron, Felix, et al.
Pubblicazione: (2026)
EpiCache: Episodic KV Cache Management for Long-Term Conversation on Resource-Constrained Environments
di: Kim, Minsoo, et al.
Pubblicazione: (2025)
di: Kim, Minsoo, et al.
Pubblicazione: (2025)
TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization
di: Yao, Dingyu, et al.
Pubblicazione: (2025)
di: Yao, Dingyu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Syntax Without Semantics: Teaching Large Language Models to Code in an Unseen Language
di: Kumar, Vinayshekhar Bannihatti, et al.
Pubblicazione: (2026) -
Neural Breadcrumbs: Membership Inference Attacks on LLMs Through Hidden State and Attention Pattern Analysis
di: Makhija, Disha, et al.
Pubblicazione: (2025) -
When Facts Change: Probing LLMs on Evolving Knowledge with evolveQA
di: Nakshatri, Nishanth Sridhar, et al.
Pubblicazione: (2025) -
Multi-Faceted Evaluation of Tool-Augmented Dialogue Systems
di: Hou, Zhaoyi Joey, et al.
Pubblicazione: (2025) -
FairGen: Controlling Sensitive Attributes for Fair Generations in Diffusion Models via Adaptive Latent Guidance
di: Kang, Mintong, et al.
Pubblicazione: (2025)