SemShareKV: Efficient KVCache Sharing for Semantically Similar Prompts via Token-Level LSH Matching
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Xinye, Mastorakis, Spyridon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Helping Large Language Models Protect Themselves: An Enhanced Filtering and Summarization System
von: Muhaimin, Sheikh Samit, et al.
Veröffentlicht: (2025)
von: Muhaimin, Sheikh Samit, et al.
Veröffentlicht: (2025)
KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing
von: Yang, Yifei, et al.
Veröffentlicht: (2024)
von: Yang, Yifei, et al.
Veröffentlicht: (2024)
SemToken: Semantic-Aware Tokenization for Efficient Long-Context Language Modeling
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache Sharing
von: Gao, Yizhao, et al.
Veröffentlicht: (2026)
von: Gao, Yizhao, et al.
Veröffentlicht: (2026)
Beyond KV Caching: Shared Attention for Efficient LLMs
von: Liao, Bingli, et al.
Veröffentlicht: (2024)
von: Liao, Bingli, et al.
Veröffentlicht: (2024)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
Krul: Efficient State Restoration for Multi-turn Conversations with Dynamic Cross-layer KV Sharing
von: Wen, Junyi, et al.
Veröffentlicht: (2025)
von: Wen, Junyi, et al.
Veröffentlicht: (2025)
ShareLoRA: Parameter Efficient and Robust Large Language Model Fine-tuning via Shared Low-Rank Adaptation
von: Song, Yurun, et al.
Veröffentlicht: (2024)
von: Song, Yurun, et al.
Veröffentlicht: (2024)
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching
von: Nguyen, Truong, et al.
Veröffentlicht: (2026)
von: Nguyen, Truong, et al.
Veröffentlicht: (2026)
DeltaKV: Residual-Based KV Cache Compression via Long-Range Similarity
von: Hao, Jitai, et al.
Veröffentlicht: (2026)
von: Hao, Jitai, et al.
Veröffentlicht: (2026)
TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection
von: Wu, Wei, et al.
Veröffentlicht: (2024)
von: Wu, Wei, et al.
Veröffentlicht: (2024)
Shared Path: Unraveling Memorization in Multilingual LLMs through Language Similarities
von: Luo, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Luo, Xiaoyu, et al.
Veröffentlicht: (2025)
CompressKV: Semantic Retrieval Heads Know What Tokens are Not Important Before Generation
von: Lin, Xiaolin, et al.
Veröffentlicht: (2025)
von: Lin, Xiaolin, et al.
Veröffentlicht: (2025)
Team Ryu's Submission to SIGMORPHON 2024 Shared Task on Subword Tokenization
von: Li, Zilong
Veröffentlicht: (2024)
von: Li, Zilong
Veröffentlicht: (2024)
BERT-LSH: Reducing Absolute Compute For Attention
von: Li, Zezheng, et al.
Veröffentlicht: (2024)
von: Li, Zezheng, et al.
Veröffentlicht: (2024)
Leveraging KV Similarity for Online Structured Pruning in LLMs
von: Lee, Jungmin, et al.
Veröffentlicht: (2025)
von: Lee, Jungmin, et al.
Veröffentlicht: (2025)
Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering
von: Mohamed, Anas, et al.
Veröffentlicht: (2025)
von: Mohamed, Anas, et al.
Veröffentlicht: (2025)
KV Prediction for Improved Time to First Token
von: Horton, Maxwell, et al.
Veröffentlicht: (2024)
von: Horton, Maxwell, et al.
Veröffentlicht: (2024)
1 Trillion Token (1TT) Platform: A Novel Framework for Efficient Data Sharing and Compensation in Large Language Models
von: Park, Chanjun, et al.
Veröffentlicht: (2024)
von: Park, Chanjun, et al.
Veröffentlicht: (2024)
Semantic Similarity Matching for Patent Documents Using Ensemble BERT-related Model and Novel Text Processing Method
von: Yu, Liqiang, et al.
Veröffentlicht: (2024)
von: Yu, Liqiang, et al.
Veröffentlicht: (2024)
PQCache: Product Quantization-based KVCache for Long Context LLM Inference
von: Zhang, Hailin, et al.
Veröffentlicht: (2024)
von: Zhang, Hailin, et al.
Veröffentlicht: (2024)
CROP: Token-Efficient Reasoning in Large Language Models via Regularized Prompt Optimization
von: Shah, Deep, et al.
Veröffentlicht: (2026)
von: Shah, Deep, et al.
Veröffentlicht: (2026)
Metric-Fair Prompting: Treating Similar Samples Similarly
von: Wang, Jing, et al.
Veröffentlicht: (2025)
von: Wang, Jing, et al.
Veröffentlicht: (2025)
Token Alignment via Character Matching for Subword Completion
von: Athiwaratkun, Ben, et al.
Veröffentlicht: (2024)
von: Athiwaratkun, Ben, et al.
Veröffentlicht: (2024)
Share Your Attention: Transformer Weight Sharing via Matrix-based Dictionary Learning
von: Zhussip, Magauiya, et al.
Veröffentlicht: (2025)
von: Zhussip, Magauiya, et al.
Veröffentlicht: (2025)
DroidSpeak: KV Cache Sharing for Cross-LLM Communication and Multi-LLM Serving
von: Liu, Yuhan, et al.
Veröffentlicht: (2024)
von: Liu, Yuhan, et al.
Veröffentlicht: (2024)
KVReviver: Reversible KV Cache Compression with Sketch-Based Token Reconstruction
von: Yuan, Aomufei, et al.
Veröffentlicht: (2025)
von: Yuan, Aomufei, et al.
Veröffentlicht: (2025)
DynSplit-KV: Dynamic Semantic Splitting for KVCache Compression in Efficient Long-Context LLM Inference
von: Ye, Jiancai, et al.
Veröffentlicht: (2026)
von: Ye, Jiancai, et al.
Veröffentlicht: (2026)
Sharing State Between Prompts and Programs
von: Cheng, Ellie Y., et al.
Veröffentlicht: (2025)
von: Cheng, Ellie Y., et al.
Veröffentlicht: (2025)
SemBench: A Universal Semantic Framework for LLM Evaluation
von: Zubillaga, Mikel, et al.
Veröffentlicht: (2026)
von: Zubillaga, Mikel, et al.
Veröffentlicht: (2026)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
von: Liu, Guangda, et al.
Veröffentlicht: (2025)
von: Liu, Guangda, et al.
Veröffentlicht: (2025)
Rethinking Word Similarity: Semantic Similarity through Classification Confusion
von: Zhou, Kaitlyn, et al.
Veröffentlicht: (2025)
von: Zhou, Kaitlyn, et al.
Veröffentlicht: (2025)
MiSS: Revisiting the Trade-off in LoRA with an Efficient Shard-Sharing Structure
von: Kang, Jiale, et al.
Veröffentlicht: (2024)
von: Kang, Jiale, et al.
Veröffentlicht: (2024)
Token-Guard: Towards Token-Level Hallucination Control via Self-Checking Decoding
von: Zhu, Yifan, et al.
Veröffentlicht: (2026)
von: Zhu, Yifan, et al.
Veröffentlicht: (2026)
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models
von: Couturier, Camille, et al.
Veröffentlicht: (2025)
von: Couturier, Camille, et al.
Veröffentlicht: (2025)
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
von: Feng, Yuan, et al.
Veröffentlicht: (2024)
von: Feng, Yuan, et al.
Veröffentlicht: (2024)
Tracing the Roots of Facts in Multilingual Language Models: Independent, Shared, and Transferred Knowledge
von: Zhao, Xin, et al.
Veröffentlicht: (2024)
von: Zhao, Xin, et al.
Veröffentlicht: (2024)
Efficient and Effective Prompt Tuning via Prompt Decomposition and Compressed Outer Product
von: Lan, Pengxiang, et al.
Veröffentlicht: (2025)
von: Lan, Pengxiang, et al.
Veröffentlicht: (2025)
IntactKV: Improving Large Language Model Quantization by Keeping Pivot Tokens Intact
von: Liu, Ruikang, et al.
Veröffentlicht: (2024)
von: Liu, Ruikang, et al.
Veröffentlicht: (2024)
Generative Caching for Structurally Similar Prompts and Responses
von: Chakraborty, Sarthak, et al.
Veröffentlicht: (2025)
von: Chakraborty, Sarthak, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Helping Large Language Models Protect Themselves: An Enhanced Filtering and Summarization System
von: Muhaimin, Sheikh Samit, et al.
Veröffentlicht: (2025) -
KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing
von: Yang, Yifei, et al.
Veröffentlicht: (2024) -
SemToken: Semantic-Aware Tokenization for Efficient Long-Context Language Modeling
von: Liu, Dong, et al.
Veröffentlicht: (2025) -
HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache Sharing
von: Gao, Yizhao, et al.
Veröffentlicht: (2026) -
Beyond KV Caching: Shared Attention for Efficient LLMs
von: Liao, Bingli, et al.
Veröffentlicht: (2024)