EvolKV: Evolutionary KV Cache Compression for LLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Bohan, Chai, Yekun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KV Cache Recycling to Expand Usable Context Capacity in Low Parameter LLMs
by: Pandey, Prashant
Published: (2025)
by: Pandey, Prashant
Published: (2025)
EvolVE: Evolutionary Search for LLM-based Verilog Generation and Optimization
by: Hsin, Wei-Po, et al.
Published: (2026)
by: Hsin, Wei-Po, et al.
Published: (2026)
Decomposing Evolutionary Mixture-of-LoRA Architectures: The Routing Lever, the Lifecycle Penalty, and a Substrate-Conditional Boundary
by: Kumaresan, Ramchand
Published: (2026)
by: Kumaresan, Ramchand
Published: (2026)
LLM-FE: Automated Feature Engineering for Tabular Data with LLMs as Evolutionary Optimizers
by: Abhyankar, Nikhil, et al.
Published: (2025)
by: Abhyankar, Nikhil, et al.
Published: (2025)
Hysteresis Activation Function for Efficient Inference
by: Kimhi, Moshe, et al.
Published: (2024)
by: Kimhi, Moshe, et al.
Published: (2024)
BrainTransformers: SNN-LLM
by: Tang, Zhengzheng, et al.
Published: (2024)
by: Tang, Zhengzheng, et al.
Published: (2024)
Multimodal LLM-assisted Evolutionary Search for Programmatic Control Policies
by: Hu, Qinglong, et al.
Published: (2025)
by: Hu, Qinglong, et al.
Published: (2025)
SpikeLLM: Scaling up Spiking Neural Network to Large Language Models via Saliency-based Spiking
by: Xing, Xingrun, et al.
Published: (2024)
by: Xing, Xingrun, et al.
Published: (2024)
Experience-Based Evolutionary Algorithms for Expensive Optimization
by: Yu, Xunzhao, et al.
Published: (2023)
by: Yu, Xunzhao, et al.
Published: (2023)
Evolutionary Multi-Objective Optimization of Large Language Model Prompts for Balancing Sentiments
by: Baumann, Jill, et al.
Published: (2024)
by: Baumann, Jill, et al.
Published: (2024)
EvoGPT-f: An Evolutionary GPT Framework for Benchmarking Formal Math Languages
by: Mercer, Johnathan
Published: (2024)
by: Mercer, Johnathan
Published: (2024)
Compute Allocation in Evolutionary Search: From Depth-Breadth to Multi-Armed Bandits
by: Xing, Sixue, et al.
Published: (2026)
by: Xing, Sixue, et al.
Published: (2026)
When Large Language Models Meet Evolutionary Algorithms: Potential Enhancements and Challenges
by: Wang, Chao, et al.
Published: (2024)
by: Wang, Chao, et al.
Published: (2024)
Evolutionary Policy Optimization
by: Mustafaoglu, Zelal Su "Lain", et al.
Published: (2025)
by: Mustafaoglu, Zelal Su "Lain", et al.
Published: (2025)
Evolutionary Multitasking AUC Optimization
by: Wang, Chao, et al.
Published: (2022)
by: Wang, Chao, et al.
Published: (2022)
Diffusion Models are Evolutionary Algorithms
by: Zhang, Yanbo, et al.
Published: (2024)
by: Zhang, Yanbo, et al.
Published: (2024)
Evolutionary Dynamic Optimization and Machine Learning
by: Boulesnane, Abdennour
Published: (2023)
by: Boulesnane, Abdennour
Published: (2023)
eVAE: Evolutionary Variational Autoencoder
by: Wu, Zhangkai, et al.
Published: (2023)
by: Wu, Zhangkai, et al.
Published: (2023)
What Makes an LLM a Good Optimizer? A Trajectory Analysis of LLM-Guided Evolutionary Search
by: Zhang, Xinhao, et al.
Published: (2026)
by: Zhang, Xinhao, et al.
Published: (2026)
Natural Evolutionary Search meets Probabilistic Numerics
by: Osselin, Pierre, et al.
Published: (2025)
by: Osselin, Pierre, et al.
Published: (2025)
Analysis and Explainability of LLMs Via Evolutionary Methods
by: Gallagher, Shannon K., et al.
Published: (2026)
by: Gallagher, Shannon K., et al.
Published: (2026)
Evolutionary Context Search for Automated Skill Acquisition
by: Sun, Qi, et al.
Published: (2026)
by: Sun, Qi, et al.
Published: (2026)
Deep Learning-Based Operators for Evolutionary Algorithms
by: Shem-Tov, Eliad, et al.
Published: (2024)
by: Shem-Tov, Eliad, et al.
Published: (2024)
TensorGPT: Efficient Compression of Large Language Models based on Tensor-Train Decomposition
by: Xu, Mingxue, et al.
Published: (2023)
by: Xu, Mingxue, et al.
Published: (2023)
Neuro-Evolutionary Approach to Physics-Aware Symbolic Regression
by: Kubalík, Jiří, et al.
Published: (2025)
by: Kubalík, Jiří, et al.
Published: (2025)
Code World Models for Parameter Control in Evolutionary Algorithms
by: Sartori, Camilo Chacón, et al.
Published: (2026)
by: Sartori, Camilo Chacón, et al.
Published: (2026)
Sharpness-Aware Minimization for Evolutionary Feature Construction in Regression
by: Zhang, Hengzhe, et al.
Published: (2024)
by: Zhang, Hengzhe, et al.
Published: (2024)
Pre-trained Language Models Learn Remarkably Accurate Representations of Numbers
by: Kadlčík, Marek, et al.
Published: (2025)
by: Kadlčík, Marek, et al.
Published: (2025)
AP-BMM: Approximating Capability-Cost Pareto Sets of LLMs via Asynchronous Prior-Guided Bayesian Model Merging
by: Chen, Kesheng, et al.
Published: (2025)
by: Chen, Kesheng, et al.
Published: (2025)
Intelligent Neural Networks: From Layered Architectures to Graph-Organized Intelligence
by: Salomon, Antoine
Published: (2025)
by: Salomon, Antoine
Published: (2025)
An In-depth Walkthrough on Evolution of Neural Machine Translation
by: Jagtap, Rohan, et al.
Published: (2020)
by: Jagtap, Rohan, et al.
Published: (2020)
Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models
by: Dong, Peijie, et al.
Published: (2024)
by: Dong, Peijie, et al.
Published: (2024)
SpikeLM: Towards General Spike-Driven Language Modeling via Elastic Bi-Spiking Mechanisms
by: Xing, Xingrun, et al.
Published: (2024)
by: Xing, Xingrun, et al.
Published: (2024)
Large Language Models Suffer From Their Own Output: An Analysis of the Self-Consuming Training Loop
by: Briesch, Martin, et al.
Published: (2023)
by: Briesch, Martin, et al.
Published: (2023)
Genetic Instruct: Scaling up Synthetic Generation of Coding Instructions for Large Language Models
by: Majumdar, Somshubra, et al.
Published: (2024)
by: Majumdar, Somshubra, et al.
Published: (2024)
Large Language Models for Tuning Evolution Strategies
by: Kramer, Oliver
Published: (2024)
by: Kramer, Oliver
Published: (2024)
On the Power of Convolution Augmented Transformer
by: Li, Mingchen, et al.
Published: (2024)
by: Li, Mingchen, et al.
Published: (2024)
A Hormone-inspired Emotion Layer for Transformer language models (HELT)
by: Reda, Eslam, et al.
Published: (2026)
by: Reda, Eslam, et al.
Published: (2026)
HAT: Hardware-Aware Transformers for Efficient Natural Language Processing
by: Wang, Hanrui, et al.
Published: (2020)
by: Wang, Hanrui, et al.
Published: (2020)
SpikingSSMs: Learning Long Sequences with Sparse and Parallel Spiking State Space Models
by: Shen, Shuaijie, et al.
Published: (2024)
by: Shen, Shuaijie, et al.
Published: (2024)
Similar Items
-
KV Cache Recycling to Expand Usable Context Capacity in Low Parameter LLMs
by: Pandey, Prashant
Published: (2025) -
EvolVE: Evolutionary Search for LLM-based Verilog Generation and Optimization
by: Hsin, Wei-Po, et al.
Published: (2026) -
Decomposing Evolutionary Mixture-of-LoRA Architectures: The Routing Lever, the Lifecycle Penalty, and a Substrate-Conditional Boundary
by: Kumaresan, Ramchand
Published: (2026) -
LLM-FE: Automated Feature Engineering for Tabular Data with LLMs as Evolutionary Optimizers
by: Abhyankar, Nikhil, et al.
Published: (2025) -
Hysteresis Activation Function for Efficient Inference
by: Kimhi, Moshe, et al.
Published: (2024)