Keyformer: KV Cache Reduction through Key Tokens Selection for Efficient Generative Inference
Fuente:
arXiv
Salvato in:
| Autori principali: | Adnan, Muhammad, Arunkumar, Akhil, Jain, Gaurav, Nair, Prashant J., Soloveychik, Ilya, Kamath, Purushotham |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LLM Inference Acceleration via Efficient Operation Fusion
di: Salmani, Mahsa, et al.
Pubblicazione: (2025)
di: Salmani, Mahsa, et al.
Pubblicazione: (2025)
Accurate Block Quantization in LLMs with Outliers
di: Trukhanov, Nikita, et al.
Pubblicazione: (2024)
di: Trukhanov, Nikita, et al.
Pubblicazione: (2024)
VeriCache: Turning Lossy KV Cache into Lossless LLM Inference
di: Yao, Jiayi, et al.
Pubblicazione: (2026)
di: Yao, Jiayi, et al.
Pubblicazione: (2026)
Comparative Characterization of KV Cache Management Strategies for LLM Inference
di: Mamo, Oteo, et al.
Pubblicazione: (2026)
di: Mamo, Oteo, et al.
Pubblicazione: (2026)
SPEC CPU: The Next Generation
di: Madhav, Mahesh, et al.
Pubblicazione: (2026)
di: Madhav, Mahesh, et al.
Pubblicazione: (2026)
Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache Management Beyond GPU Limits
di: Kim, Dowon, et al.
Pubblicazione: (2025)
di: Kim, Dowon, et al.
Pubblicazione: (2025)
Titanus: Enabling KV Cache Pruning and Quantization On-the-Fly for LLM Acceleration
di: Chen, Peilin, et al.
Pubblicazione: (2025)
di: Chen, Peilin, et al.
Pubblicazione: (2025)
UniCAIM: A Unified CAM/CIM Architecture with Static-Dynamic KV Cache Pruning for Efficient Long-Context LLM Inference
di: Xu, Weikai, et al.
Pubblicazione: (2025)
di: Xu, Weikai, et al.
Pubblicazione: (2025)
Unconventional Universal Computation in Babbage's Analytical Engine
di: Rojas, Raul
Pubblicazione: (2024)
di: Rojas, Raul
Pubblicazione: (2024)
Heterogeneous Acceleration Pipeline for Recommendation System Training
di: Adnan, Muhammad, et al.
Pubblicazione: (2022)
di: Adnan, Muhammad, et al.
Pubblicazione: (2022)
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator
di: Wang, Zhican, et al.
Pubblicazione: (2025)
di: Wang, Zhican, et al.
Pubblicazione: (2025)
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
di: Fang, Yunhua, et al.
Pubblicazione: (2025)
di: Fang, Yunhua, et al.
Pubblicazione: (2025)
Workload-Aware Hardware Accelerator Mining for Distributed Deep Learning Training
di: Adnan, Muhammad, et al.
Pubblicazione: (2024)
di: Adnan, Muhammad, et al.
Pubblicazione: (2024)
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
di: Ganjihal, Sanjeev Rao
Pubblicazione: (2026)
di: Ganjihal, Sanjeev Rao
Pubblicazione: (2026)
The Dirty Secret of SSDs: Embodied Carbon
di: Tannu, Swamit, et al.
Pubblicazione: (2022)
di: Tannu, Swamit, et al.
Pubblicazione: (2022)
Formalising CXL Cache Coherence
di: Tan, Chengsong, et al.
Pubblicazione: (2024)
di: Tan, Chengsong, et al.
Pubblicazione: (2024)
TokenStack: A Heterogeneous HBM-PIM Architecture and Runtime for Efficient LLM Inference
di: Li, Zhuoran, et al.
Pubblicazione: (2026)
di: Li, Zhuoran, et al.
Pubblicazione: (2026)
Systolic Arrays and Structured Pruning Co-design for Efficient Transformers in Edge Systems
di: Palacios, Pedro, et al.
Pubblicazione: (2024)
di: Palacios, Pedro, et al.
Pubblicazione: (2024)
Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
di: Kim, Minsu, et al.
Pubblicazione: (2025)
di: Kim, Minsu, et al.
Pubblicazione: (2025)
HillInfer: Efficient Long-Context LLM Inference on the Edge with Hierarchical KV Eviction using SmartSSD
di: Sun, He, et al.
Pubblicazione: (2026)
di: Sun, He, et al.
Pubblicazione: (2026)
DiSC: Resolution-Scalable Acceleration of Diffusion Models by Exploiting Sparsity and Cached Token Reuse with Hash-based Distribution
di: Yoon, Jieon, et al.
Pubblicazione: (2026)
di: Yoon, Jieon, et al.
Pubblicazione: (2026)
CGRA4ML: A Hardware/Software Framework to Implement Neural Networks for Scientific Edge Computing
di: Abarajithan, G, et al.
Pubblicazione: (2024)
di: Abarajithan, G, et al.
Pubblicazione: (2024)
Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge Computing
di: Xia, Tianhua, et al.
Pubblicazione: (2025)
di: Xia, Tianhua, et al.
Pubblicazione: (2025)
SATA: Sparsity-Aware Scheduling for Selective Token Attention
di: Fan, Zhenkun, et al.
Pubblicazione: (2026)
di: Fan, Zhenkun, et al.
Pubblicazione: (2026)
Finite-Time Lyapunov Exponent Calculation on FPGA using High-Level Synthesis Tools
di: de Castro, Manuel, et al.
Pubblicazione: (2024)
di: de Castro, Manuel, et al.
Pubblicazione: (2024)
HG-PIPE: Vision Transformer Acceleration with Hybrid-Grained Pipeline
di: Guo, Qingyu, et al.
Pubblicazione: (2024)
di: Guo, Qingyu, et al.
Pubblicazione: (2024)
DCI: A Coordinated Allocation and Filling Workload-Aware Dual-Cache Allocation GNN Inference Acceleration System
di: Luo, Yi, et al.
Pubblicazione: (2025)
di: Luo, Yi, et al.
Pubblicazione: (2025)
Understanding Inference-Time Token Allocation and Coverage Limits in Agentic Hardware Verification
di: Patel, Vihaan, et al.
Pubblicazione: (2026)
di: Patel, Vihaan, et al.
Pubblicazione: (2026)
SliceMoE: Bit-Sliced Expert Caching under Miss-Rate Constraints for Efficient MoE Inference
di: Choi, Yuseon, et al.
Pubblicazione: (2025)
di: Choi, Yuseon, et al.
Pubblicazione: (2025)
PiKV: KV Cache Management System for Mixture of Experts
di: Liu, Dong, et al.
Pubblicazione: (2025)
di: Liu, Dong, et al.
Pubblicazione: (2025)
Ouroboros: Wafer-Scale SRAM CIM with Token-Grained Pipelining for Large Language Model Inference
di: Liu, Yiqi, et al.
Pubblicazione: (2026)
di: Liu, Yiqi, et al.
Pubblicazione: (2026)
3RSeT: Read Disturbance Rate Reduction in STT-MRAM Caches by Selective Tag Comparison
di: Cheshmikhani, Elham, et al.
Pubblicazione: (2025)
di: Cheshmikhani, Elham, et al.
Pubblicazione: (2025)
Commercial Evaluation of Zero-Skipping MAC Design for Bit Sparsity Exploitation in DL Inference
di: Nair, Harideep, et al.
Pubblicazione: (2024)
di: Nair, Harideep, et al.
Pubblicazione: (2024)
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
di: li, Fei, et al.
Pubblicazione: (2026)
di: li, Fei, et al.
Pubblicazione: (2026)
PIVOT- Input-aware Path Selection for Energy-efficient ViT Inference
di: Moitra, Abhishek, et al.
Pubblicazione: (2024)
di: Moitra, Abhishek, et al.
Pubblicazione: (2024)
Exploring DRAM Cache Prefetching for Pooled Memory
di: Tirumalasetty, Chandrahas, et al.
Pubblicazione: (2024)
di: Tirumalasetty, Chandrahas, et al.
Pubblicazione: (2024)
Potential and Limitation of High-Frequency Cores and Caches
di: Pai, Kunal, et al.
Pubblicazione: (2024)
di: Pai, Kunal, et al.
Pubblicazione: (2024)
TDRAM: Tag-enhanced DRAM for Efficient Caching
di: Babaie, Maryam, et al.
Pubblicazione: (2024)
di: Babaie, Maryam, et al.
Pubblicazione: (2024)
Bit-Flip Vulnerability of Shared KV-Cache Blocks in LLM Serving Systems
di: Yamamoto, Yuji, et al.
Pubblicazione: (2026)
di: Yamamoto, Yuji, et al.
Pubblicazione: (2026)
Adaptive Multi-Objective Tiered Storage Configuration for KV Cache in LLM Service
di: Zheng, Xianzhe, et al.
Pubblicazione: (2026)
di: Zheng, Xianzhe, et al.
Pubblicazione: (2026)
Documenti analoghi
-
LLM Inference Acceleration via Efficient Operation Fusion
di: Salmani, Mahsa, et al.
Pubblicazione: (2025) -
Accurate Block Quantization in LLMs with Outliers
di: Trukhanov, Nikita, et al.
Pubblicazione: (2024) -
VeriCache: Turning Lossy KV Cache into Lossless LLM Inference
di: Yao, Jiayi, et al.
Pubblicazione: (2026) -
Comparative Characterization of KV Cache Management Strategies for LLM Inference
di: Mamo, Oteo, et al.
Pubblicazione: (2026) -
SPEC CPU: The Next Generation
di: Madhav, Mahesh, et al.
Pubblicazione: (2026)