KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Jingbo, Hou, Bairu, Wei, Wei, Bao, Yujia, Chang, Shiyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Ares: Adaptive Reasoning Effort Selection for Efficient LLM Agents
von: Yang, Jingbo, et al.
Veröffentlicht: (2026)
von: Yang, Jingbo, et al.
Veröffentlicht: (2026)
A Probabilistic Framework for LLM Hallucination Detection via Belief Tree Propagation
von: Hou, Bairu, et al.
Veröffentlicht: (2024)
von: Hou, Bairu, et al.
Veröffentlicht: (2024)
WebDART: Dynamic Decomposition and Re-planning for Complex Web Tasks
von: Yang, Jingbo, et al.
Veröffentlicht: (2025)
von: Yang, Jingbo, et al.
Veröffentlicht: (2025)
Decomposing Uncertainty for Large Language Models through Input Clarification Ensembling
von: Hou, Bairu, et al.
Veröffentlicht: (2023)
von: Hou, Bairu, et al.
Veröffentlicht: (2023)
EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction
von: Ji, Shiyu, et al.
Veröffentlicht: (2026)
von: Ji, Shiyu, et al.
Veröffentlicht: (2026)
Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing
von: Ji, Jiabao, et al.
Veröffentlicht: (2024)
von: Ji, Jiabao, et al.
Veröffentlicht: (2024)
EFIM: Efficient Serving of LLMs for Infilling Tasks with Improved KV Cache Reuse
von: Guo, Tianyu, et al.
Veröffentlicht: (2025)
von: Guo, Tianyu, et al.
Veröffentlicht: (2025)
Instruction-Following Pruning for Large Language Models
von: Hou, Bairu, et al.
Veröffentlicht: (2025)
von: Hou, Bairu, et al.
Veröffentlicht: (2025)
Layer-Condensed KV Cache for Efficient Inference of Large Language Models
von: Wu, Haoyi, et al.
Veröffentlicht: (2024)
von: Wu, Haoyi, et al.
Veröffentlicht: (2024)
FlashDLM: Accelerating Diffusion Language Model Inference via Efficient KV Caching and Guided Diffusion
von: Hu, Zhanqiu, et al.
Veröffentlicht: (2025)
von: Hu, Zhanqiu, et al.
Veröffentlicht: (2025)
KVShare: An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse
von: Yang, Huan, et al.
Veröffentlicht: (2025)
von: Yang, Huan, et al.
Veröffentlicht: (2025)
A Method for Building Large Language Models with Predefined KV Cache Capacity
von: Yi, Zhonghua, et al.
Veröffentlicht: (2024)
von: Yi, Zhonghua, et al.
Veröffentlicht: (2024)
CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems
von: Yang, Jingbo, et al.
Veröffentlicht: (2026)
von: Yang, Jingbo, et al.
Veröffentlicht: (2026)
dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching
von: Liu, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Liu, Zhiyuan, et al.
Veröffentlicht: (2025)
DASH-KV: Accelerating Long-Context LLM Inference via Asymmetric KV Cache Hashing
von: Guo, Jinyu, et al.
Veröffentlicht: (2026)
von: Guo, Jinyu, et al.
Veröffentlicht: (2026)
ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning
von: Hou, Bairu, et al.
Veröffentlicht: (2025)
von: Hou, Bairu, et al.
Veröffentlicht: (2025)
dKV-Cache: The Cache for Diffusion Language Models
von: Ma, Xinyin, et al.
Veröffentlicht: (2025)
von: Ma, Xinyin, et al.
Veröffentlicht: (2025)
Advancing the Robustness of Large Language Models through Self-Denoised Smoothing
von: Ji, Jiabao, et al.
Veröffentlicht: (2024)
von: Ji, Jiabao, et al.
Veröffentlicht: (2024)
Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Query
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
AhaKV: Adaptive Holistic Attention-Driven KV Cache Eviction for Efficient Inference of Large Language Models
von: Gu, Yifeng, et al.
Veröffentlicht: (2025)
von: Gu, Yifeng, et al.
Veröffentlicht: (2025)
LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models
von: Shi, Dachuan, et al.
Veröffentlicht: (2025)
von: Shi, Dachuan, et al.
Veröffentlicht: (2025)
WindowKV: Task-Adaptive Group-Wise KV Cache Window Selection for Efficient LLM Inference
von: Zuo, Youhui, et al.
Veröffentlicht: (2025)
von: Zuo, Youhui, et al.
Veröffentlicht: (2025)
GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2025)
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2025)
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
von: Liu, Akide, et al.
Veröffentlicht: (2024)
von: Liu, Akide, et al.
Veröffentlicht: (2024)
SemantiCache: Efficient KV Cache Compression via Semantic Chunking and Clustered Merging
von: Wu, Shunlong, et al.
Veröffentlicht: (2026)
von: Wu, Shunlong, et al.
Veröffentlicht: (2026)
VecInfer: Efficient LLM Inference with Low-Bit KV Cache via Outlier-Suppressed Vector Quantization
von: Yao, Dingyu, et al.
Veröffentlicht: (2025)
von: Yao, Dingyu, et al.
Veröffentlicht: (2025)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization
von: Yao, Dingyu, et al.
Veröffentlicht: (2025)
von: Yao, Dingyu, et al.
Veröffentlicht: (2025)
HyperRAG: Enhancing Quality-Efficiency Tradeoffs in Retrieval-Augmented Generation with Reranker KV-Cache Reuse
von: An, Yuwei, et al.
Veröffentlicht: (2025)
von: An, Yuwei, et al.
Veröffentlicht: (2025)
xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction
von: Chang, Chi-Chih, et al.
Veröffentlicht: (2025)
von: Chang, Chi-Chih, et al.
Veröffentlicht: (2025)
RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression
von: Behnam, Payman, et al.
Veröffentlicht: (2025)
von: Behnam, Payman, et al.
Veröffentlicht: (2025)
IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
von: Bai, Yushi, et al.
Veröffentlicht: (2026)
von: Bai, Yushi, et al.
Veröffentlicht: (2026)
QAQ: Quality Adaptive Quantization for LLM KV Cache
von: Dong, Shichen, et al.
Veröffentlicht: (2024)
von: Dong, Shichen, et al.
Veröffentlicht: (2024)
EntropyCache: Decoded Token Entropy Guided KV Caching for Diffusion Language Models
von: Cheong, Minsoo, et al.
Veröffentlicht: (2026)
von: Cheong, Minsoo, et al.
Veröffentlicht: (2026)
Efficient Long-Context LLM Inference via KV Cache Clustering
von: Hu, Jie, et al.
Veröffentlicht: (2025)
von: Hu, Jie, et al.
Veröffentlicht: (2025)
In-context KV-Cache Eviction for LLMs via Attention-Gate
von: Zeng, Zihao, et al.
Veröffentlicht: (2024)
von: Zeng, Zihao, et al.
Veröffentlicht: (2024)
VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference Acceleration
von: Tu, Dezhan, et al.
Veröffentlicht: (2024)
von: Tu, Dezhan, et al.
Veröffentlicht: (2024)
AnTKV: Anchor Token-Aware Sub-Bit Vector Quantization for KV Cache in Large Language Models
von: Li, Zeyu, et al.
Veröffentlicht: (2025)
von: Li, Zeyu, et al.
Veröffentlicht: (2025)
PromptBridge: Cross-Model Prompt Transfer for Large Language Models
von: Wang, Yaxuan, et al.
Veröffentlicht: (2025)
von: Wang, Yaxuan, et al.
Veröffentlicht: (2025)
ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Ares: Adaptive Reasoning Effort Selection for Efficient LLM Agents
von: Yang, Jingbo, et al.
Veröffentlicht: (2026) -
A Probabilistic Framework for LLM Hallucination Detection via Belief Tree Propagation
von: Hou, Bairu, et al.
Veröffentlicht: (2024) -
WebDART: Dynamic Decomposition and Re-planning for Complex Web Tasks
von: Yang, Jingbo, et al.
Veröffentlicht: (2025) -
Decomposing Uncertainty for Large Language Models through Input Clarification Ensembling
von: Hou, Bairu, et al.
Veröffentlicht: (2023) -
EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction
von: Ji, Shiyu, et al.
Veröffentlicht: (2026)