SemShareKV: Efficient KVCache Sharing for Semantically Similar Prompts via Token-Level LSH Matching
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Xinye, Mastorakis, Spyridon |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Helping Large Language Models Protect Themselves: An Enhanced Filtering and Summarization System
by: Muhaimin, Sheikh Samit, et al.
Published: (2025)
by: Muhaimin, Sheikh Samit, et al.
Published: (2025)
KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing
by: Yang, Yifei, et al.
Published: (2024)
by: Yang, Yifei, et al.
Published: (2024)
SemToken: Semantic-Aware Tokenization for Efficient Long-Context Language Modeling
by: Liu, Dong, et al.
Published: (2025)
by: Liu, Dong, et al.
Published: (2025)
HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache Sharing
by: Gao, Yizhao, et al.
Published: (2026)
by: Gao, Yizhao, et al.
Published: (2026)
Beyond KV Caching: Shared Attention for Efficient LLMs
by: Liao, Bingli, et al.
Published: (2024)
by: Liao, Bingli, et al.
Published: (2024)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
by: Zhu, Yuxuan, et al.
Published: (2025)
by: Zhu, Yuxuan, et al.
Published: (2025)
Krul: Efficient State Restoration for Multi-turn Conversations with Dynamic Cross-layer KV Sharing
by: Wen, Junyi, et al.
Published: (2025)
by: Wen, Junyi, et al.
Published: (2025)
ShareLoRA: Parameter Efficient and Robust Large Language Model Fine-tuning via Shared Low-Rank Adaptation
by: Song, Yurun, et al.
Published: (2024)
by: Song, Yurun, et al.
Published: (2024)
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching
by: Nguyen, Truong, et al.
Published: (2026)
by: Nguyen, Truong, et al.
Published: (2026)
DeltaKV: Residual-Based KV Cache Compression via Long-Range Similarity
by: Hao, Jitai, et al.
Published: (2026)
by: Hao, Jitai, et al.
Published: (2026)
TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection
by: Wu, Wei, et al.
Published: (2024)
by: Wu, Wei, et al.
Published: (2024)
Shared Path: Unraveling Memorization in Multilingual LLMs through Language Similarities
by: Luo, Xiaoyu, et al.
Published: (2025)
by: Luo, Xiaoyu, et al.
Published: (2025)
CompressKV: Semantic Retrieval Heads Know What Tokens are Not Important Before Generation
by: Lin, Xiaolin, et al.
Published: (2025)
by: Lin, Xiaolin, et al.
Published: (2025)
Team Ryu's Submission to SIGMORPHON 2024 Shared Task on Subword Tokenization
by: Li, Zilong
Published: (2024)
by: Li, Zilong
Published: (2024)
BERT-LSH: Reducing Absolute Compute For Attention
by: Li, Zezheng, et al.
Published: (2024)
by: Li, Zezheng, et al.
Published: (2024)
Leveraging KV Similarity for Online Structured Pruning in LLMs
by: Lee, Jungmin, et al.
Published: (2025)
by: Lee, Jungmin, et al.
Published: (2025)
Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering
by: Mohamed, Anas, et al.
Published: (2025)
by: Mohamed, Anas, et al.
Published: (2025)
KV Prediction for Improved Time to First Token
by: Horton, Maxwell, et al.
Published: (2024)
by: Horton, Maxwell, et al.
Published: (2024)
1 Trillion Token (1TT) Platform: A Novel Framework for Efficient Data Sharing and Compensation in Large Language Models
by: Park, Chanjun, et al.
Published: (2024)
by: Park, Chanjun, et al.
Published: (2024)
Semantic Similarity Matching for Patent Documents Using Ensemble BERT-related Model and Novel Text Processing Method
by: Yu, Liqiang, et al.
Published: (2024)
by: Yu, Liqiang, et al.
Published: (2024)
PQCache: Product Quantization-based KVCache for Long Context LLM Inference
by: Zhang, Hailin, et al.
Published: (2024)
by: Zhang, Hailin, et al.
Published: (2024)
CROP: Token-Efficient Reasoning in Large Language Models via Regularized Prompt Optimization
by: Shah, Deep, et al.
Published: (2026)
by: Shah, Deep, et al.
Published: (2026)
Metric-Fair Prompting: Treating Similar Samples Similarly
by: Wang, Jing, et al.
Published: (2025)
by: Wang, Jing, et al.
Published: (2025)
Token Alignment via Character Matching for Subword Completion
by: Athiwaratkun, Ben, et al.
Published: (2024)
by: Athiwaratkun, Ben, et al.
Published: (2024)
Share Your Attention: Transformer Weight Sharing via Matrix-based Dictionary Learning
by: Zhussip, Magauiya, et al.
Published: (2025)
by: Zhussip, Magauiya, et al.
Published: (2025)
DroidSpeak: KV Cache Sharing for Cross-LLM Communication and Multi-LLM Serving
by: Liu, Yuhan, et al.
Published: (2024)
by: Liu, Yuhan, et al.
Published: (2024)
KVReviver: Reversible KV Cache Compression with Sketch-Based Token Reconstruction
by: Yuan, Aomufei, et al.
Published: (2025)
by: Yuan, Aomufei, et al.
Published: (2025)
DynSplit-KV: Dynamic Semantic Splitting for KVCache Compression in Efficient Long-Context LLM Inference
by: Ye, Jiancai, et al.
Published: (2026)
by: Ye, Jiancai, et al.
Published: (2026)
Sharing State Between Prompts and Programs
by: Cheng, Ellie Y., et al.
Published: (2025)
by: Cheng, Ellie Y., et al.
Published: (2025)
SemBench: A Universal Semantic Framework for LLM Evaluation
by: Zubillaga, Mikel, et al.
Published: (2026)
by: Zubillaga, Mikel, et al.
Published: (2026)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
by: Liu, Guangda, et al.
Published: (2025)
by: Liu, Guangda, et al.
Published: (2025)
Rethinking Word Similarity: Semantic Similarity through Classification Confusion
by: Zhou, Kaitlyn, et al.
Published: (2025)
by: Zhou, Kaitlyn, et al.
Published: (2025)
MiSS: Revisiting the Trade-off in LoRA with an Efficient Shard-Sharing Structure
by: Kang, Jiale, et al.
Published: (2024)
by: Kang, Jiale, et al.
Published: (2024)
Token-Guard: Towards Token-Level Hallucination Control via Self-Checking Decoding
by: Zhu, Yifan, et al.
Published: (2026)
by: Zhu, Yifan, et al.
Published: (2026)
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models
by: Couturier, Camille, et al.
Published: (2025)
by: Couturier, Camille, et al.
Published: (2025)
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
by: Feng, Yuan, et al.
Published: (2024)
by: Feng, Yuan, et al.
Published: (2024)
Tracing the Roots of Facts in Multilingual Language Models: Independent, Shared, and Transferred Knowledge
by: Zhao, Xin, et al.
Published: (2024)
by: Zhao, Xin, et al.
Published: (2024)
Efficient and Effective Prompt Tuning via Prompt Decomposition and Compressed Outer Product
by: Lan, Pengxiang, et al.
Published: (2025)
by: Lan, Pengxiang, et al.
Published: (2025)
IntactKV: Improving Large Language Model Quantization by Keeping Pivot Tokens Intact
by: Liu, Ruikang, et al.
Published: (2024)
by: Liu, Ruikang, et al.
Published: (2024)
Generative Caching for Structurally Similar Prompts and Responses
by: Chakraborty, Sarthak, et al.
Published: (2025)
by: Chakraborty, Sarthak, et al.
Published: (2025)
Similar Items
-
Helping Large Language Models Protect Themselves: An Enhanced Filtering and Summarization System
by: Muhaimin, Sheikh Samit, et al.
Published: (2025) -
KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing
by: Yang, Yifei, et al.
Published: (2024) -
SemToken: Semantic-Aware Tokenization for Efficient Long-Context Language Modeling
by: Liu, Dong, et al.
Published: (2025) -
HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache Sharing
by: Gao, Yizhao, et al.
Published: (2026) -
Beyond KV Caching: Shared Attention for Efficient LLMs
by: Liao, Bingli, et al.
Published: (2024)