Efficient LLM Inference with Kcache
Fuente:
arXiv
Saved in:
| Main Authors: | He, Qiaozhi, Wu, Zhihua |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring Scaling Laws for Local SGD in Large Language Model Training
by: He, Qiaozhi, et al.
Published: (2024)
by: He, Qiaozhi, et al.
Published: (2024)
ChuXin: 1.6B Technical Report
by: Zhuang, Xiaomin, et al.
Published: (2024)
by: Zhuang, Xiaomin, et al.
Published: (2024)
Code Comparison Tuning for Code Large Language Models
by: Jiang, Yufan, et al.
Published: (2024)
by: Jiang, Yufan, et al.
Published: (2024)
RecycleGPT: An Autoregressive Language Model with Recyclable Module
by: Jiang, Yufan, et al.
Published: (2023)
by: Jiang, Yufan, et al.
Published: (2023)
A Systematic Study of Cross-Layer KV Sharing for Efficient LLM Inference
by: Wu, You, et al.
Published: (2024)
by: Wu, You, et al.
Published: (2024)
YOCO++: Enhancing YOCO with KV Residual Connections for Efficient LLM Inference
by: Wu, You, et al.
Published: (2026)
by: Wu, You, et al.
Published: (2026)
Efficient Long-Context LLM Inference via KV Cache Clustering
by: Hu, Jie, et al.
Published: (2025)
by: Hu, Jie, et al.
Published: (2025)
CHAI: Clustered Head Attention for Efficient LLM Inference
by: Agarwal, Saurabh, et al.
Published: (2024)
by: Agarwal, Saurabh, et al.
Published: (2024)
READER: Retrieval-Assisted Drafter for Efficient LLM Inference
by: Divilkovskiy, Maxim, et al.
Published: (2025)
by: Divilkovskiy, Maxim, et al.
Published: (2025)
Tutorial Proposal: Speculative Decoding for Efficient LLM Inference
by: Xia, Heming, et al.
Published: (2025)
by: Xia, Heming, et al.
Published: (2025)
Universal Model Routing for Efficient LLM Inference
by: Jitkrittum, Wittawat, et al.
Published: (2025)
by: Jitkrittum, Wittawat, et al.
Published: (2025)
FLRC: Fine-grained Low-Rank Compressor for Efficient LLM Inference
by: Lu, Yu-Chen, et al.
Published: (2025)
by: Lu, Yu-Chen, et al.
Published: (2025)
Optimizing Transformer based on high-performance optimizer for predicting employment sentiment in American social media content
by: Wang, Feiyang, et al.
Published: (2024)
by: Wang, Feiyang, et al.
Published: (2024)
DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inference
by: Liu, Xiang, et al.
Published: (2025)
by: Liu, Xiang, et al.
Published: (2025)
ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
by: Liang, Yesheng, et al.
Published: (2025)
by: Liang, Yesheng, et al.
Published: (2025)
BUZZ: Beehive-structured Sparse KV Cache with Segmented Heavy Hitters for Efficient LLM Inference
by: Zhao, Junqi, et al.
Published: (2024)
by: Zhao, Junqi, et al.
Published: (2024)
Progressive Mixed-Precision Decoding for Efficient LLM Inference
by: Chen, Hao Mark, et al.
Published: (2024)
by: Chen, Hao Mark, et al.
Published: (2024)
Probe and Skip: Self-Predictive Token Skipping for Efficient Long-Context LLM Inference
by: Wu, Zimeng, et al.
Published: (2026)
by: Wu, Zimeng, et al.
Published: (2026)
RevMUX: Data Multiplexing with Reversible Adapters for Efficient LLM Batch Inference
by: Xu, Yige, et al.
Published: (2024)
by: Xu, Yige, et al.
Published: (2024)
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
by: Xiao, Guangxuan, et al.
Published: (2024)
by: Xiao, Guangxuan, et al.
Published: (2024)
LoRA-Drop: Temporal LoRA Decoding for Efficient LLM Inference
by: Rajabzadeh, Hossein, et al.
Published: (2026)
by: Rajabzadeh, Hossein, et al.
Published: (2026)
Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length
by: Ma, Xuezhe, et al.
Published: (2024)
by: Ma, Xuezhe, et al.
Published: (2024)
Mask Tokens as Prophet: Fine-Grained Cache Eviction for Efficient dLLM Inference
by: Huang, Jianuo, et al.
Published: (2025)
by: Huang, Jianuo, et al.
Published: (2025)
Creative Convergence or Imitation? Genre-Specific Homogeneity in LLM-Generated Chinese Literature
by: Ma, Yuanchi, et al.
Published: (2026)
by: Ma, Yuanchi, et al.
Published: (2026)
Noise Supervised Contrastive Learning and Feature-Perturbed for Anomalous Sound Detection
by: Huang, Shun, et al.
Published: (2025)
by: Huang, Shun, et al.
Published: (2025)
Spoken Language Identification with Pre-trained Models and Margin Loss
by: Fang, Zhihua, et al.
Published: (2026)
by: Fang, Zhihua, et al.
Published: (2026)
Efficient Inference for Large Reasoning Models: A Survey
by: Liu, Yue, et al.
Published: (2025)
by: Liu, Yue, et al.
Published: (2025)
Efficient Reasoning via Thought-Training and Thought-Free Inference
by: Wu, Canhui, et al.
Published: (2025)
by: Wu, Canhui, et al.
Published: (2025)
Layer-Condensed KV Cache for Efficient Inference of Large Language Models
by: Wu, Haoyi, et al.
Published: (2024)
by: Wu, Haoyi, et al.
Published: (2024)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
by: Liu, Xiang, et al.
Published: (2025)
by: Liu, Xiang, et al.
Published: (2025)
XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference
by: Monteiro, João, et al.
Published: (2024)
by: Monteiro, João, et al.
Published: (2024)
Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
by: Tang, Jiaming, et al.
Published: (2024)
by: Tang, Jiaming, et al.
Published: (2024)
Fast KVzip: Efficient and Accurate LLM Inference with Gated KV Eviction
by: Kim, Jang-Hyun, et al.
Published: (2026)
by: Kim, Jang-Hyun, et al.
Published: (2026)
AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM Inference
by: He, Zhuomin, et al.
Published: (2025)
by: He, Zhuomin, et al.
Published: (2025)
Accelerating Diffusion LLM Inference via Local Determinism Propagation
by: Kong, Fanheng, et al.
Published: (2025)
by: Kong, Fanheng, et al.
Published: (2025)
LLM Inference Acceleration via Efficient Operation Fusion
by: Salmani, Mahsa, et al.
Published: (2025)
by: Salmani, Mahsa, et al.
Published: (2025)
Enhancing Multi-Agent Consensus through Third-Party LLM Integration: Analyzing Uncertainty and Mitigating Hallucinations in Large Language Models
by: Duan, Zhihua, et al.
Published: (2024)
by: Duan, Zhihua, et al.
Published: (2024)
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
by: Fu, Qichen, et al.
Published: (2024)
by: Fu, Qichen, et al.
Published: (2024)
WindowKV: Task-Adaptive Group-Wise KV Cache Window Selection for Efficient LLM Inference
by: Zuo, Youhui, et al.
Published: (2025)
by: Zuo, Youhui, et al.
Published: (2025)
NeuronMLP: Efficient LLM Inference via Singular Value Decomposition Compression and Tiling on AWS Trainium
by: Song, Dinghong, et al.
Published: (2025)
by: Song, Dinghong, et al.
Published: (2025)
Similar Items
-
Exploring Scaling Laws for Local SGD in Large Language Model Training
by: He, Qiaozhi, et al.
Published: (2024) -
ChuXin: 1.6B Technical Report
by: Zhuang, Xiaomin, et al.
Published: (2024) -
Code Comparison Tuning for Code Large Language Models
by: Jiang, Yufan, et al.
Published: (2024) -
RecycleGPT: An Autoregressive Language Model with Recyclable Module
by: Jiang, Yufan, et al.
Published: (2023) -
A Systematic Study of Cross-Layer KV Sharing for Efficient LLM Inference
by: Wu, You, et al.
Published: (2024)