KV Cache Recycling to Expand Usable Context Capacity in Low Parameter LLMs
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Pandey, Prashant |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Lossless Compression of Neural Network Components: Weights, Checkpoints, and K/V Caches in Low-Precision Formats
von: Heilper, Anat, et al.
Veröffentlicht: (2025)
von: Heilper, Anat, et al.
Veröffentlicht: (2025)
Polysemanticity and Capacity in Neural Networks
von: Scherlis, Adam, et al.
Veröffentlicht: (2022)
von: Scherlis, Adam, et al.
Veröffentlicht: (2022)
Parameter-Efficient Fine-Tuning of LLMs with Mixture of Space Experts
von: Zhang, Buze, et al.
Veröffentlicht: (2026)
von: Zhang, Buze, et al.
Veröffentlicht: (2026)
Measuring the Measures: Discriminative Capacity of Representational Similarity Metrics Across Model Families
von: Wu, Jialin, et al.
Veröffentlicht: (2025)
von: Wu, Jialin, et al.
Veröffentlicht: (2025)
EvolKV: Evolutionary KV Cache Compression for LLM Inference
von: Yu, Bohan, et al.
Veröffentlicht: (2025)
von: Yu, Bohan, et al.
Veröffentlicht: (2025)
Reinforced In-Context Black-Box Optimization
von: Song, Lei, et al.
Veröffentlicht: (2024)
von: Song, Lei, et al.
Veröffentlicht: (2024)
Provably Optimal Memory Capacity for Modern Hopfield Models: Transformer-Compatible Dense Associative Memories as Spherical Codes
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2024)
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2024)
Globally Optimal Training of Spiking Neural Networks via Parameter Reconstruction
von: Udupi, Himanshu, et al.
Veröffentlicht: (2026)
von: Udupi, Himanshu, et al.
Veröffentlicht: (2026)
General-Purpose In-Context Learning by Meta-Learning Transformers
von: Kirsch, Louis, et al.
Veröffentlicht: (2022)
von: Kirsch, Louis, et al.
Veröffentlicht: (2022)
QL-LSTM: A Parameter-Efficient LSTM for Stable Long-Sequence Modeling
von: Nti, Isaac Kofi
Veröffentlicht: (2025)
von: Nti, Isaac Kofi
Veröffentlicht: (2025)
Adapting Rule Representation With Four-Parameter Beta Distribution for Learning Classifier Systems
von: Shiraishi, Hiroki, et al.
Veröffentlicht: (2025)
von: Shiraishi, Hiroki, et al.
Veröffentlicht: (2025)
LLM-Meta-SR: In-Context Learning for Evolving Selection Operators in Symbolic Regression
von: Zhang, Hengzhe, et al.
Veröffentlicht: (2025)
von: Zhang, Hengzhe, et al.
Veröffentlicht: (2025)
Context-sensitive neocortical neurons transform the effectiveness and efficiency of neural information processing
von: Ahmed, Khubaib, et al.
Veröffentlicht: (2022)
von: Ahmed, Khubaib, et al.
Veröffentlicht: (2022)
Neuro-Symbolic Activation Discovery: Transferring Mathematical Structures from Physics to Ecology for Parameter-Efficient Neural Networks
von: Hajbi, Anas
Veröffentlicht: (2026)
von: Hajbi, Anas
Veröffentlicht: (2026)
Enabling Robust In-Context Memory and Rapid Task Adaptation in Transformers with Hebbian and Gradient-Based Plasticity
von: Chaudhary, Siddharth
Veröffentlicht: (2025)
von: Chaudhary, Siddharth
Veröffentlicht: (2025)
Optuna vs Code Llama: Are LLMs a New Paradigm for Hyperparameter Tuning?
von: Kochnev, Roman, et al.
Veröffentlicht: (2025)
von: Kochnev, Roman, et al.
Veröffentlicht: (2025)
GigaEvo: An Open Source Optimization Framework Powered By LLMs And Evolution Algorithms
von: Khrulkov, Valentin, et al.
Veröffentlicht: (2025)
von: Khrulkov, Valentin, et al.
Veröffentlicht: (2025)
SAN: Hypothesizing Long-Term Synaptic Development and Neural Engram Mechanism in Scalable Model's Parameter-Efficient Fine-Tuning
von: Dai, Gaole, et al.
Veröffentlicht: (2024)
von: Dai, Gaole, et al.
Veröffentlicht: (2024)
LANCE: Low Rank Activation Compression for Efficient On-Device Continual Learning
von: Apolinario, Marco Paul E., et al.
Veröffentlicht: (2025)
von: Apolinario, Marco Paul E., et al.
Veröffentlicht: (2025)
A Low Latency Adaptive Coding Spiking Framework for Deep Reinforcement Learning
von: Qin, Lang, et al.
Veröffentlicht: (2022)
von: Qin, Lang, et al.
Veröffentlicht: (2022)
Gradient-Free Training of Spiking Neural Networks via Low-Rank Evolution Strategies
von: Patankar, Dhruv, et al.
Veröffentlicht: (2026)
von: Patankar, Dhruv, et al.
Veröffentlicht: (2026)
Graph Learning-based Regional Heavy Rainfall Prediction Using Low-Cost Rain Gauges
von: Salcedo, Edwin
Veröffentlicht: (2024)
von: Salcedo, Edwin
Veröffentlicht: (2024)
Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention
von: Munkhdalai, Tsendsuren, et al.
Veröffentlicht: (2024)
von: Munkhdalai, Tsendsuren, et al.
Veröffentlicht: (2024)
A Methodology to Study the Impact of Spiking Neural Network Parameters considering Event-Based Automotive Data
von: Bano, Iqra, et al.
Veröffentlicht: (2024)
von: Bano, Iqra, et al.
Veröffentlicht: (2024)
The Impact of Structural Changes on Learning Capacity in the Fly Olfactory Neural Circuit
von: Xie, Katherine, et al.
Veröffentlicht: (2025)
von: Xie, Katherine, et al.
Veröffentlicht: (2025)
Jailbreaking LLMs' Safeguard with Universal Magic Words for Text Embedding Models
von: Liang, Haoyu, et al.
Veröffentlicht: (2025)
von: Liang, Haoyu, et al.
Veröffentlicht: (2025)
NOBLE: Accelerating Transformers with Nonlinear Low-Rank Branches
von: Smith, Ethan
Veröffentlicht: (2026)
von: Smith, Ethan
Veröffentlicht: (2026)
LLM-FE: Automated Feature Engineering for Tabular Data with LLMs as Evolutionary Optimizers
von: Abhyankar, Nikhil, et al.
Veröffentlicht: (2025)
von: Abhyankar, Nikhil, et al.
Veröffentlicht: (2025)
Straight to Zero: Why Linearly Decaying the Learning Rate to Zero Works Best for LLMs
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)
DL101 Neural Network Outputs and Loss Functions
von: Berzal, Fernando
Veröffentlicht: (2025)
von: Berzal, Fernando
Veröffentlicht: (2025)
Evolved Sample Weights for Bias Mitigation: Effectiveness Depends on the Fairness Objective
von: Saini, Anil K., et al.
Veröffentlicht: (2025)
von: Saini, Anil K., et al.
Veröffentlicht: (2025)
Efficient ANN-SNN Conversion with Error Compensation Learning
von: Liu, Chang, et al.
Veröffentlicht: (2025)
von: Liu, Chang, et al.
Veröffentlicht: (2025)
Enhancing Cross Entropy with a Linearly Adaptive Loss Function for Optimized Classification Performance
von: Shim, Jae Wan
Veröffentlicht: (2025)
von: Shim, Jae Wan
Veröffentlicht: (2025)
Particle swarm optimization for online sparse streaming feature selection under uncertainty
von: Xu, Ruiyang
Veröffentlicht: (2025)
von: Xu, Ruiyang
Veröffentlicht: (2025)
Empirical Investigation into Configuring Echo State Networks for Representative Benchmark Problem Domains
von: Weborg, Brooke R., et al.
Veröffentlicht: (2025)
von: Weborg, Brooke R., et al.
Veröffentlicht: (2025)
HyperGraphX: Graph Transductive Learning with Hyperdimensional Computing and Message Passing
von: Cong, Guojing, et al.
Veröffentlicht: (2025)
von: Cong, Guojing, et al.
Veröffentlicht: (2025)
EOE: Evolutionary Optimization of Experts for Training Language Models
von: Chen, Yingshi
Veröffentlicht: (2025)
von: Chen, Yingshi
Veröffentlicht: (2025)
Node Preservation and its Effect on Crossover in Cartesian Genetic Programming
von: Kocherovsky, Mark, et al.
Veröffentlicht: (2025)
von: Kocherovsky, Mark, et al.
Veröffentlicht: (2025)
Understanding Transformer Optimization via Gradient Heterogeneity
von: Tomihari, Akiyoshi, et al.
Veröffentlicht: (2025)
von: Tomihari, Akiyoshi, et al.
Veröffentlicht: (2025)
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning
von: Qiu, Xin, et al.
Veröffentlicht: (2025)
von: Qiu, Xin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Lossless Compression of Neural Network Components: Weights, Checkpoints, and K/V Caches in Low-Precision Formats
von: Heilper, Anat, et al.
Veröffentlicht: (2025) -
Polysemanticity and Capacity in Neural Networks
von: Scherlis, Adam, et al.
Veröffentlicht: (2022) -
Parameter-Efficient Fine-Tuning of LLMs with Mixture of Space Experts
von: Zhang, Buze, et al.
Veröffentlicht: (2026) -
Measuring the Measures: Discriminative Capacity of Representational Similarity Metrics Across Model Families
von: Wu, Jialin, et al.
Veröffentlicht: (2025) -
EvolKV: Evolutionary KV Cache Compression for LLM Inference
von: Yu, Bohan, et al.
Veröffentlicht: (2025)