On the Efficacy of Eviction Policy for Key-Value Constrained Generative Language Model Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ren, Siyu, Zhu, Kenny Q. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AhaKV: Adaptive Holistic Attention-Driven KV Cache Eviction for Efficient Inference of Large Language Models
von: Gu, Yifeng, et al.
Veröffentlicht: (2025)
von: Gu, Yifeng, et al.
Veröffentlicht: (2025)
Reformulating KV Cache Eviction Problem for Long-Context LLM Inference
von: Mai, Tho, et al.
Veröffentlicht: (2026)
von: Mai, Tho, et al.
Veröffentlicht: (2026)
IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference
von: Yang, Xintong, et al.
Veröffentlicht: (2026)
von: Yang, Xintong, et al.
Veröffentlicht: (2026)
Stepwise Alignment for Constrained Language Model Policy Optimization
von: Wachi, Akifumi, et al.
Veröffentlicht: (2024)
von: Wachi, Akifumi, et al.
Veröffentlicht: (2024)
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
von: Feng, Yuan, et al.
Veröffentlicht: (2024)
von: Feng, Yuan, et al.
Veröffentlicht: (2024)
SynthEHR-Eviction: Enhancing Eviction SDoH Detection with LLM-Augmented Synthetic EHR Data
von: Yao, Zonghai, et al.
Veröffentlicht: (2025)
von: Yao, Zonghai, et al.
Veröffentlicht: (2025)
EMO: Earth Mover Distance Optimization for Auto-Regressive Language Modeling
von: Ren, Siyu, et al.
Veröffentlicht: (2023)
von: Ren, Siyu, et al.
Veröffentlicht: (2023)
Audit-of-Understanding: Posterior-Constrained Inference for Mathematical Reasoning in Language Models
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
von: Abdaljalil, Samir, et al.
Veröffentlicht: (2025)
Eliciting Trustworthiness Priors of Large Language Models via Economic Games
von: Yan, Siyu, et al.
Veröffentlicht: (2026)
von: Yan, Siyu, et al.
Veröffentlicht: (2026)
Judge Q: Trainable Queries for Optimized Information Retention in KV Cache Eviction
von: Liu, Yijun, et al.
Veröffentlicht: (2025)
von: Liu, Yijun, et al.
Veröffentlicht: (2025)
Measuring Human and AI Values Based on Generative Psychometrics with Large Language Models
von: Ye, Haoran, et al.
Veröffentlicht: (2024)
von: Ye, Haoran, et al.
Veröffentlicht: (2024)
Inference-Time Language Model Alignment via Integrated Value Guidance
von: Liu, Zhixuan, et al.
Veröffentlicht: (2024)
von: Liu, Zhixuan, et al.
Veröffentlicht: (2024)
A Cognitive Evaluation Benchmark of Image Reasoning and Description for Large Vision-Language Models
von: Song, Xiujie, et al.
Veröffentlicht: (2024)
von: Song, Xiujie, et al.
Veröffentlicht: (2024)
Generative Psycho-Lexical Approach for Constructing Value Systems in Large Language Models
von: Ye, Haoran, et al.
Veröffentlicht: (2025)
von: Ye, Haoran, et al.
Veröffentlicht: (2025)
Topic Over Source: The Key to Effective Data Mixing for Language Models Pre-training
von: Peng, Jiahui, et al.
Veröffentlicht: (2025)
von: Peng, Jiahui, et al.
Veröffentlicht: (2025)
Activation-aware Probe-Query: Effective Key-Value Retrieval for Long-Context LLMs Inference
von: Xiao, Qingfa, et al.
Veröffentlicht: (2025)
von: Xiao, Qingfa, et al.
Veröffentlicht: (2025)
KV-Runahead: Scalable Causal LLM Inference by Parallel Key-Value Cache Generation
von: Cho, Minsik, et al.
Veröffentlicht: (2024)
von: Cho, Minsik, et al.
Veröffentlicht: (2024)
Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Query
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
RobustKV: Defending Large Language Models against Jailbreak Attacks via KV Eviction
von: Jiang, Tanqiu, et al.
Veröffentlicht: (2024)
von: Jiang, Tanqiu, et al.
Veröffentlicht: (2024)
Representations as Language: An Information-Theoretic Framework for Interpretability
von: Conklin, Henry, et al.
Veröffentlicht: (2024)
von: Conklin, Henry, et al.
Veröffentlicht: (2024)
MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference
von: Li, Kunxi, et al.
Veröffentlicht: (2025)
von: Li, Kunxi, et al.
Veröffentlicht: (2025)
G-KV: Decoding-Time KV Cache Eviction with Global Attention
von: Liao, Mengqi, et al.
Veröffentlicht: (2025)
von: Liao, Mengqi, et al.
Veröffentlicht: (2025)
Schema Key Wording as an Instruction Channel in Structured Generation under Constrained Decoding
von: Le, Yifan
Veröffentlicht: (2026)
von: Le, Yifan
Veröffentlicht: (2026)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
von: Wang, Guangtao, et al.
Veröffentlicht: (2025)
von: Wang, Guangtao, et al.
Veröffentlicht: (2025)
Key-Point-Driven Mathematical Reasoning Distillation of Large Language Model
von: Zhu, Xunyu, et al.
Veröffentlicht: (2024)
von: Zhu, Xunyu, et al.
Veröffentlicht: (2024)
Translate-and-Revise: Boosting Large Language Models for Constrained Translation
von: Huang, Pengcheng, et al.
Veröffentlicht: (2024)
von: Huang, Pengcheng, et al.
Veröffentlicht: (2024)
Large Language Model for Patent Concept Generation
von: Ren, Runtao, et al.
Veröffentlicht: (2024)
von: Ren, Runtao, et al.
Veröffentlicht: (2024)
WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More
von: Yue, Yuxuan, et al.
Veröffentlicht: (2024)
von: Yue, Yuxuan, et al.
Veröffentlicht: (2024)
Conversational Control with Ontologies for Large Language Models: A Lightweight Framework for Constrained Generation
von: Gendron, Barbara, et al.
Veröffentlicht: (2026)
von: Gendron, Barbara, et al.
Veröffentlicht: (2026)
Unlocking Anticipatory Text Generation: A Constrained Approach for Large Language Models Decoding
von: Tu, Lifu, et al.
Veröffentlicht: (2023)
von: Tu, Lifu, et al.
Veröffentlicht: (2023)
Boosting Self-Efficacy and Performance of Large Language Models via Verbal Efficacy Stimulations
von: Chen, Rui, et al.
Veröffentlicht: (2025)
von: Chen, Rui, et al.
Veröffentlicht: (2025)
Benchmarking Multi-National Value Alignment for Large Language Models
von: Shi, Weijie, et al.
Veröffentlicht: (2025)
von: Shi, Weijie, et al.
Veröffentlicht: (2025)
Abductive Inference in Retrieval-Augmented Language Models: Generating and Validating Missing Premises
von: Lin, Shiyin
Veröffentlicht: (2025)
von: Lin, Shiyin
Veröffentlicht: (2025)
CoLLEGe: Concept Embedding Generation for Large Language Models
von: Teehan, Ryan, et al.
Veröffentlicht: (2024)
von: Teehan, Ryan, et al.
Veröffentlicht: (2024)
Parallel Key-Value Cache Fusion for Position Invariant RAG
von: Oh, Philhoon, et al.
Veröffentlicht: (2025)
von: Oh, Philhoon, et al.
Veröffentlicht: (2025)
TaskBench: Benchmarking Large Language Models for Task Automation
von: Shen, Yongliang, et al.
Veröffentlicht: (2023)
von: Shen, Yongliang, et al.
Veröffentlicht: (2023)
Scaling Parameter-Constrained Language Models with Quality Data
von: Chang, Ernie, et al.
Veröffentlicht: (2024)
von: Chang, Ernie, et al.
Veröffentlicht: (2024)
Large Language Models for Constrained-Based Causal Discovery
von: Cohrs, Kai-Hendrik, et al.
Veröffentlicht: (2024)
von: Cohrs, Kai-Hendrik, et al.
Veröffentlicht: (2024)
Causal Order: The Key to Leveraging Imperfect Experts in Causal Inference
von: Vashishtha, Aniket, et al.
Veröffentlicht: (2023)
von: Vashishtha, Aniket, et al.
Veröffentlicht: (2023)
Symbol-LLM: Towards Foundational Symbol-centric Interface For Large Language Models
von: Xu, Fangzhi, et al.
Veröffentlicht: (2023)
von: Xu, Fangzhi, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
AhaKV: Adaptive Holistic Attention-Driven KV Cache Eviction for Efficient Inference of Large Language Models
von: Gu, Yifeng, et al.
Veröffentlicht: (2025) -
Reformulating KV Cache Eviction Problem for Long-Context LLM Inference
von: Mai, Tho, et al.
Veröffentlicht: (2026) -
IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference
von: Yang, Xintong, et al.
Veröffentlicht: (2026) -
Stepwise Alignment for Constrained Language Model Policy Optimization
von: Wachi, Akifumi, et al.
Veröffentlicht: (2024) -
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
von: Feng, Yuan, et al.
Veröffentlicht: (2024)