Knowledge Packs: Zero-Token Knowledge Delivery via KV Cache Injection
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Pustovit, Andrey |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Accurate KV Cache Quantization with Outlier Tokens Tracing
von: Su, Yi, et al.
Veröffentlicht: (2025)
von: Su, Yi, et al.
Veröffentlicht: (2025)
EntropyCache: Decoded Token Entropy Guided KV Caching for Diffusion Language Models
von: Cheong, Minsoo, et al.
Veröffentlicht: (2026)
von: Cheong, Minsoo, et al.
Veröffentlicht: (2026)
CTkvr: KV Cache Retrieval for Long-Context LLMs via Centroid then Token Indexing
von: Lu, Kuan, et al.
Veröffentlicht: (2025)
von: Lu, Kuan, et al.
Veröffentlicht: (2025)
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression
von: Zhang, Jiebin, et al.
Veröffentlicht: (2024)
von: Zhang, Jiebin, et al.
Veröffentlicht: (2024)
KVReviver: Reversible KV Cache Compression with Sketch-Based Token Reconstruction
von: Yuan, Aomufei, et al.
Veröffentlicht: (2025)
von: Yuan, Aomufei, et al.
Veröffentlicht: (2025)
KV Cache Offloading for Context-Intensive Tasks
von: Bocharnikov, Andrey, et al.
Veröffentlicht: (2026)
von: Bocharnikov, Andrey, et al.
Veröffentlicht: (2026)
EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction
von: Ji, Shiyu, et al.
Veröffentlicht: (2026)
von: Ji, Shiyu, et al.
Veröffentlicht: (2026)
LagKV: Lag-Relative Information of the KV Cache Tells Which Tokens Are Important
von: Liang, Manlai, et al.
Veröffentlicht: (2025)
von: Liang, Manlai, et al.
Veröffentlicht: (2025)
CAOTE: KV Cache Selection for LLMs via Attention Output Error-Based Token Eviction
von: Goel, Raghavv, et al.
Veröffentlicht: (2025)
von: Goel, Raghavv, et al.
Veröffentlicht: (2025)
Beyond RAG: Task-Aware KV Cache Compression for Comprehensive Knowledge Reasoning
von: Corallo, Giulio, et al.
Veröffentlicht: (2025)
von: Corallo, Giulio, et al.
Veröffentlicht: (2025)
GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2025)
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2025)
dKV-Cache: The Cache for Diffusion Language Models
von: Ma, Xinyin, et al.
Veröffentlicht: (2025)
von: Ma, Xinyin, et al.
Veröffentlicht: (2025)
SemantiCache: Efficient KV Cache Compression via Semantic Chunking and Clustered Merging
von: Wu, Shunlong, et al.
Veröffentlicht: (2026)
von: Wu, Shunlong, et al.
Veröffentlicht: (2026)
DASH-KV: Accelerating Long-Context LLM Inference via Asymmetric KV Cache Hashing
von: Guo, Jinyu, et al.
Veröffentlicht: (2026)
von: Guo, Jinyu, et al.
Veröffentlicht: (2026)
Task-KV: Task-aware KV Cache Optimization via Semantic Differentiation of Attention Heads
von: He, Xingyang, et al.
Veröffentlicht: (2025)
von: He, Xingyang, et al.
Veröffentlicht: (2025)
KORE: Enhancing Knowledge Injection for Large Multimodal Models via Knowledge-Oriented Controls
von: Jiang, Kailin, et al.
Veröffentlicht: (2025)
von: Jiang, Kailin, et al.
Veröffentlicht: (2025)
TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization
von: Yao, Dingyu, et al.
Veröffentlicht: (2025)
von: Yao, Dingyu, et al.
Veröffentlicht: (2025)
DeltaKV: Residual-Based KV Cache Compression via Long-Range Similarity
von: Hao, Jitai, et al.
Veröffentlicht: (2026)
von: Hao, Jitai, et al.
Veröffentlicht: (2026)
TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection
von: Wu, Wei, et al.
Veröffentlicht: (2024)
von: Wu, Wei, et al.
Veröffentlicht: (2024)
NestedKV: Nested Memory Routing for Long-Context KV Cache Compression
von: Chen, Hong, et al.
Veröffentlicht: (2026)
von: Chen, Hong, et al.
Veröffentlicht: (2026)
CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective
von: Feng, Yuan, et al.
Veröffentlicht: (2025)
von: Feng, Yuan, et al.
Veröffentlicht: (2025)
RefreshKV: Updating Small KV Cache During Long-form Generation
von: Xu, Fangyuan, et al.
Veröffentlicht: (2024)
von: Xu, Fangyuan, et al.
Veröffentlicht: (2024)
ZSMerge: Zero-Shot KV Cache Compression for Memory-Efficient Long-Context LLMs
von: Liu, Xin, et al.
Veröffentlicht: (2025)
von: Liu, Xin, et al.
Veröffentlicht: (2025)
Efficient Long-Context LLM Inference via KV Cache Clustering
von: Hu, Jie, et al.
Veröffentlicht: (2025)
von: Hu, Jie, et al.
Veröffentlicht: (2025)
AnTKV: Anchor Token-Aware Sub-Bit Vector Quantization for KV Cache in Large Language Models
von: Li, Zeyu, et al.
Veröffentlicht: (2025)
von: Li, Zeyu, et al.
Veröffentlicht: (2025)
CaliDrop: KV Cache Compression with Calibration
von: Su, Yi, et al.
Veröffentlicht: (2025)
von: Su, Yi, et al.
Veröffentlicht: (2025)
xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction
von: Chang, Chi-Chih, et al.
Veröffentlicht: (2025)
von: Chang, Chi-Chih, et al.
Veröffentlicht: (2025)
EvolKV: Evolutionary KV Cache Compression for LLM Inference
von: Yu, Bohan, et al.
Veröffentlicht: (2025)
von: Yu, Bohan, et al.
Veröffentlicht: (2025)
One Size Does Not Fit All: Token-Wise Adaptive Compression for KV Cache
von: Lu, Liming, et al.
Veröffentlicht: (2026)
von: Lu, Liming, et al.
Veröffentlicht: (2026)
Decoupling Reasoning and Knowledge Injection for In-Context Knowledge Editing
von: Wang, Changyue, et al.
Veröffentlicht: (2025)
von: Wang, Changyue, et al.
Veröffentlicht: (2025)
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs
von: Zhou, Xiabin, et al.
Veröffentlicht: (2024)
von: Zhou, Xiabin, et al.
Veröffentlicht: (2024)
GraphKV: Breaking the Static Selection Paradigm with Graph-Based KV Cache Eviction
von: Li, Xuelin, et al.
Veröffentlicht: (2025)
von: Li, Xuelin, et al.
Veröffentlicht: (2025)
In-context KV-Cache Eviction for LLMs via Attention-Gate
von: Zeng, Zihao, et al.
Veröffentlicht: (2024)
von: Zeng, Zihao, et al.
Veröffentlicht: (2024)
Lossless KV Cache Compression to 2%
von: Yang, Zhen, et al.
Veröffentlicht: (2024)
von: Yang, Zhen, et al.
Veröffentlicht: (2024)
G-KV: Decoding-Time KV Cache Eviction with Global Attention
von: Liao, Mengqi, et al.
Veröffentlicht: (2025)
von: Liao, Mengqi, et al.
Veröffentlicht: (2025)
R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Query
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
Efficient Knowledge Injection in LLMs via Self-Distillation
von: Kujanpää, Kalle, et al.
Veröffentlicht: (2024)
von: Kujanpää, Kalle, et al.
Veröffentlicht: (2024)
LKV: End-to-End Learning of Head-wise Budgets and Token Selection for LLM KV Cache Eviction
von: Zhou, Enshuai, et al.
Veröffentlicht: (2026)
von: Zhou, Enshuai, et al.
Veröffentlicht: (2026)
HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache Sharing
von: Gao, Yizhao, et al.
Veröffentlicht: (2026)
von: Gao, Yizhao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Accurate KV Cache Quantization with Outlier Tokens Tracing
von: Su, Yi, et al.
Veröffentlicht: (2025) -
EntropyCache: Decoded Token Entropy Guided KV Caching for Diffusion Language Models
von: Cheong, Minsoo, et al.
Veröffentlicht: (2026) -
CTkvr: KV Cache Retrieval for Long-Context LLMs via Centroid then Token Indexing
von: Lu, Kuan, et al.
Veröffentlicht: (2025) -
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression
von: Zhang, Jiebin, et al.
Veröffentlicht: (2024) -
KVReviver: Reversible KV Cache Compression with Sketch-Based Token Reconstruction
von: Yuan, Aomufei, et al.
Veröffentlicht: (2025)