Finch: Prompt-guided Key-Value Cache Compression
Fuente:
arXiv
Saved in:
| Main Authors: | Corallo, Giulio, Papotti, Paolo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Parallel Context-of-Experts Decoding for Retrieval Augmented Generation
by: Corallo, Giulio, et al.
Published: (2026)
by: Corallo, Giulio, et al.
Published: (2026)
Beyond RAG: Task-Aware KV Cache Compression for Comprehensive Knowledge Reasoning
by: Corallo, Giulio, et al.
Published: (2025)
by: Corallo, Giulio, et al.
Published: (2025)
GoldFinch: High Performance RWKV/Transformer Hybrid with Linear Pre-Fill and Extreme KV-Cache Compression
by: Goldstein, Daniel, et al.
Published: (2024)
by: Goldstein, Daniel, et al.
Published: (2024)
Rethinking Key-Value Cache Compression Techniques for Large Language Model Serving
by: Gao, Wei, et al.
Published: (2025)
by: Gao, Wei, et al.
Published: (2025)
Latent Abstractions in Generative Diffusion Models
by: Franzese, Giulio, et al.
Published: (2024)
by: Franzese, Giulio, et al.
Published: (2024)
Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence
by: Peng, Bo, et al.
Published: (2024)
by: Peng, Bo, et al.
Published: (2024)
Constraint Decay: The Fragility of LLM Agents in Backend Code Generation
by: Dente, Francesco, et al.
Published: (2026)
by: Dente, Francesco, et al.
Published: (2026)
Parallel Key-Value Cache Fusion for Position Invariant RAG
by: Oh, Philhoon, et al.
Published: (2025)
by: Oh, Philhoon, et al.
Published: (2025)
Prompt-Based Value Steering of Large Language Models
by: Abbo, Giulio Antonio, et al.
Published: (2025)
by: Abbo, Giulio Antonio, et al.
Published: (2025)
RelationalFactQA: A Benchmark for Evaluating Tabular Fact Retrieval from Large Language Models
by: Satriani, Dario, et al.
Published: (2025)
by: Satriani, Dario, et al.
Published: (2025)
Thin Keys, Full Values: Reducing KV Cache via Low-Dimensional Attention Selection
by: Yao, Hengshuai, et al.
Published: (2026)
by: Yao, Hengshuai, et al.
Published: (2026)
The Pitfalls of KV Cache Compression
by: Chen, Alex, et al.
Published: (2025)
by: Chen, Alex, et al.
Published: (2025)
Key-Value Pair-Free Continual Learner via Task-Specific Prompt-Prototype
by: Luo, Haihua, et al.
Published: (2026)
by: Luo, Haihua, et al.
Published: (2026)
WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More
by: Yue, Yuxuan, et al.
Published: (2024)
by: Yue, Yuxuan, et al.
Published: (2024)
Lossless KV Cache Compression to 2%
by: Yang, Zhen, et al.
Published: (2024)
by: Yang, Zhen, et al.
Published: (2024)
FreqKV: Key-Value Compression in Frequency Domain for Context Window Extension
by: Kai, Jushi, et al.
Published: (2025)
by: Kai, Jushi, et al.
Published: (2025)
Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference
by: Zhu, Yue, et al.
Published: (2025)
by: Zhu, Yue, et al.
Published: (2025)
SVDq: 1.25-bit and 410x Key Cache Compression for LLM Attention
by: Yankun, Hong, et al.
Published: (2025)
by: Yankun, Hong, et al.
Published: (2025)
KV-Runahead: Scalable Causal LLM Inference by Parallel Key-Value Cache Generation
by: Cho, Minsik, et al.
Published: (2024)
by: Cho, Minsik, et al.
Published: (2024)
Combating Misinformation in the Arab World: Challenges & Opportunities
by: Abouzied, Azza, et al.
Published: (2025)
by: Abouzied, Azza, et al.
Published: (2025)
An LLM-Based Approach for Insight Generation in Data Analysis
by: Pérez, Alberto Sánchez, et al.
Published: (2025)
by: Pérez, Alberto Sánchez, et al.
Published: (2025)
When Does Value-Aware KV Eviction Help? A Fixed-Contract Diagnostic for Non-Monotone Cache Compression
by: Zhang, Ruijie, et al.
Published: (2026)
by: Zhang, Ruijie, et al.
Published: (2026)
Generative Caching for Structurally Similar Prompts and Responses
by: Chakraborty, Sarthak, et al.
Published: (2025)
by: Chakraborty, Sarthak, et al.
Published: (2025)
KVSculpt: KV Cache Compression as Distillation
by: Jiang, Bo, et al.
Published: (2026)
by: Jiang, Bo, et al.
Published: (2026)
Adaptive KV-Cache Compression without Manually Setting Budget
by: Tang, Chenxia, et al.
Published: (2025)
by: Tang, Chenxia, et al.
Published: (2025)
Palu: Compressing KV-Cache with Low-Rank Projection
by: Chang, Chi-Chih, et al.
Published: (2024)
by: Chang, Chi-Chih, et al.
Published: (2024)
Hurwitz Quaternion Multiplicative Quantization for KV Cache Compression
by: Swain, Kabir, et al.
Published: (2026)
by: Swain, Kabir, et al.
Published: (2026)
TableCache: Primary Foreign Key Guided KV Cache Precomputation for Low Latency Text-to-SQL
by: Su, Jinbo, et al.
Published: (2026)
by: Su, Jinbo, et al.
Published: (2026)
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
by: Liu, Akide, et al.
Published: (2024)
by: Liu, Akide, et al.
Published: (2024)
KeyDiff: Key Similarity-Based KV Cache Eviction for Long-Context LLM Inference in Resource-Constrained Environments
by: Park, Junyoung, et al.
Published: (2025)
by: Park, Junyoung, et al.
Published: (2025)
ThinK: Thinner Key Cache by Query-Driven Pruning
by: Xu, Yuhui, et al.
Published: (2024)
by: Xu, Yuhui, et al.
Published: (2024)
CommVQ: Commutative Vector Quantization for KV Cache Compression
by: Li, Junyan, et al.
Published: (2025)
by: Li, Junyan, et al.
Published: (2025)
HeteroCache: A Dynamic Retrieval Approach to Heterogeneous KV Cache Compression for Long-Context LLM Inference
by: Shi, Zhiyuan, et al.
Published: (2026)
by: Shi, Zhiyuan, et al.
Published: (2026)
Moment-KV: Momentum-Based Decode-Time KV Cache Compression for Long Generation
by: Jana, Soumyadeep, et al.
Published: (2026)
by: Jana, Soumyadeep, et al.
Published: (2026)
Meta-Soft: Leveraging Composable Meta-Tokens for Context-Preserving KV Cache Compression
by: Luo, Wei, et al.
Published: (2026)
by: Luo, Wei, et al.
Published: (2026)
Quantized Keys Steal Attention: Bias Correction for KV-Cache Compression in Video Diffusion
by: Tuncer, Tuna, et al.
Published: (2026)
by: Tuncer, Tuna, et al.
Published: (2026)
Prompt Cache: Modular Attention Reuse for Low-Latency Inference
by: Gim, In, et al.
Published: (2023)
by: Gim, In, et al.
Published: (2023)
CompactPrompt: A Unified Pipeline for Prompt Data Compression in LLM Workflows
by: Choi, Joong Ho, et al.
Published: (2025)
by: Choi, Joong Ho, et al.
Published: (2025)
From Similarity to Vulnerability: Key Collision Attack on LLM Semantic Caching
by: Zhang, Zhixiang, et al.
Published: (2026)
by: Zhang, Zhixiang, et al.
Published: (2026)
How Much Cache Does Reasoning Need? Depth-Cache Tradeoffs in KV-Compressed Transformers
by: Wang, Xiao
Published: (2026)
by: Wang, Xiao
Published: (2026)
Similar Items
-
Parallel Context-of-Experts Decoding for Retrieval Augmented Generation
by: Corallo, Giulio, et al.
Published: (2026) -
Beyond RAG: Task-Aware KV Cache Compression for Comprehensive Knowledge Reasoning
by: Corallo, Giulio, et al.
Published: (2025) -
GoldFinch: High Performance RWKV/Transformer Hybrid with Linear Pre-Fill and Extreme KV-Cache Compression
by: Goldstein, Daniel, et al.
Published: (2024) -
Rethinking Key-Value Cache Compression Techniques for Large Language Model Serving
by: Gao, Wei, et al.
Published: (2025) -
Latent Abstractions in Generative Diffusion Models
by: Franzese, Giulio, et al.
Published: (2024)