KVCrush: Key value cache size-reduction using similarity in head-behaviour
Fuente:
arXiv
Saved in:
| Main Authors: | Jha, Gopi Krishna, Gobriel, Sameh, Talamanova, Liubov, Jain, Nilesh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mem-Rec: Memory Efficient Recommendation System using Alternative Representation
by: Jha, Gopi Krishna, et al.
Published: (2023)
by: Jha, Gopi Krishna, et al.
Published: (2023)
TokenButler: Token Importance is Predictable
by: Akhauri, Yash, et al.
Published: (2025)
by: Akhauri, Yash, et al.
Published: (2025)
Shears: Unstructured Sparsity with Neural Low-rank Adapter Search
by: Muñoz, J. Pablo, et al.
Published: (2024)
by: Muñoz, J. Pablo, et al.
Published: (2024)
SQFT: Low-cost Model Adaptation in Low-precision Sparse Foundation Models
by: Muñoz, Juan Pablo, et al.
Published: (2024)
by: Muñoz, Juan Pablo, et al.
Published: (2024)
Low-Rank Adapters Meet Neural Architecture Search for LLM Compression
by: Muñoz, J. Pablo, et al.
Published: (2025)
by: Muñoz, J. Pablo, et al.
Published: (2025)
Automated Literature Review Using NLP Techniques and LLM-Based Retrieval-Augmented Generation
by: Ali, Nurshat Fateh, et al.
Published: (2024)
by: Ali, Nurshat Fateh, et al.
Published: (2024)
Mamba-Shedder: Post-Transformer Compression for Efficient Selective Structured State Space Models
by: Muñoz, J. Pablo, et al.
Published: (2025)
by: Muñoz, J. Pablo, et al.
Published: (2025)
Solo Connection: A Parameter Efficient Fine-Tuning Technique for Transformers
by: Pathak, Harsh Nilesh, et al.
Published: (2025)
by: Pathak, Harsh Nilesh, et al.
Published: (2025)
Exploring Design Choices for Building Language-Specific LLMs
by: Tejaswi, Atula, et al.
Published: (2024)
by: Tejaswi, Atula, et al.
Published: (2024)
Explicit Multi-head Attention for Inter-head Interaction in Large Language Models
by: Peng, Runyu, et al.
Published: (2026)
by: Peng, Runyu, et al.
Published: (2026)
On Mitigating Code LLM Hallucinations with API Documentation
by: Jain, Nihal, et al.
Published: (2024)
by: Jain, Nihal, et al.
Published: (2024)
ARGO: An Auto-Tuning Runtime System for Scalable GNN Training on Multi-Core Processor
by: Lin, Yi-Chien, et al.
Published: (2024)
by: Lin, Yi-Chien, et al.
Published: (2024)
Representation of perceived prosodic similarity of conversational feedback
by: Qian, Livia, et al.
Published: (2025)
by: Qian, Livia, et al.
Published: (2025)
LLM generation novelty through the lens of semantic similarity
by: Davydov, Philipp, et al.
Published: (2025)
by: Davydov, Philipp, et al.
Published: (2025)
Hymba: A Hybrid-head Architecture for Small Language Models
by: Dong, Xin, et al.
Published: (2024)
by: Dong, Xin, et al.
Published: (2024)
Out-of-distribution generalization via composition: a lens through induction heads in Transformers
by: Song, Jiajun, et al.
Published: (2024)
by: Song, Jiajun, et al.
Published: (2024)
Oh! We Freeze: Improving Quantized Knowledge Distillation via Signal Propagation Analysis for Large Language Models
by: Bhardwaj, Kartikeya, et al.
Published: (2024)
by: Bhardwaj, Kartikeya, et al.
Published: (2024)
Event-Keyed Summarization
by: Gantt, William, et al.
Published: (2024)
by: Gantt, William, et al.
Published: (2024)
Reasoning Towards Fairness: Mitigating Bias in Language Models through Reasoning-Guided Fine-Tuning
by: Kabra, Sanchit, et al.
Published: (2025)
by: Kabra, Sanchit, et al.
Published: (2025)
SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs
by: AbouElhamayed, Ahmed F., et al.
Published: (2025)
by: AbouElhamayed, Ahmed F., et al.
Published: (2025)
On the Emergence of Thinking in LLMs I: Searching for the Right Intuition
by: Ye, Guanghao, et al.
Published: (2025)
by: Ye, Guanghao, et al.
Published: (2025)
RAG-Modulo: Solving Sequential Tasks using Experience, Critics, and Language Models
by: Jain, Abhinav, et al.
Published: (2024)
by: Jain, Abhinav, et al.
Published: (2024)
Consistency Is the Key: Detecting Hallucinations in LLM Generated Text By Checking Inconsistencies About Key Facts
by: Gupta, Raavi, et al.
Published: (2025)
by: Gupta, Raavi, et al.
Published: (2025)
Semantic Anchors in In-Context Learning: Why Small LLMs Cannot Flip Their Labels
by: Kumar, Anantha Padmanaban Krishna
Published: (2025)
by: Kumar, Anantha Padmanaban Krishna
Published: (2025)
Dissecting Persona-Driven Reasoning in Language Models via Activation Patching
by: Poonia, Ansh, et al.
Published: (2025)
by: Poonia, Ansh, et al.
Published: (2025)
Local Prompt Optimization
by: Jain, Yash, et al.
Published: (2025)
by: Jain, Yash, et al.
Published: (2025)
Learning to Reason with Mixture of Tokens
by: Jain, Adit, et al.
Published: (2025)
by: Jain, Adit, et al.
Published: (2025)
Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models
by: Sim, Shamus, et al.
Published: (2024)
by: Sim, Shamus, et al.
Published: (2024)
Agribot: agriculture-specific question answer system
by: Jain, Naman, et al.
Published: (2025)
by: Jain, Naman, et al.
Published: (2025)
Is Training Data Quality or Quantity More Impactful to Small Language Model Performance?
by: Sajith, Aryan, et al.
Published: (2024)
by: Sajith, Aryan, et al.
Published: (2024)
One-Pass to Reason: Token Duplication and Block-Sparse Mask for Efficient Fine-Tuning on Multi-Turn Reasoning
by: Goru, Ritesh, et al.
Published: (2025)
by: Goru, Ritesh, et al.
Published: (2025)
Deep Knowledge-Infusion For Explainable Depression Detection
by: Dalal, Sumit, et al.
Published: (2024)
by: Dalal, Sumit, et al.
Published: (2024)
Hallucination is Inevitable: An Innate Limitation of Large Language Models
by: Xu, Ziwei, et al.
Published: (2024)
by: Xu, Ziwei, et al.
Published: (2024)
Are Small Language Models Ready to Compete with Large Language Models for Practical Applications?
by: Sinha, Neelabh, et al.
Published: (2024)
by: Sinha, Neelabh, et al.
Published: (2024)
Scalable Bayesian Low-Rank Adaptation of Large Language Models via Stochastic Variational Subspace Inference
by: Samplawski, Colin, et al.
Published: (2025)
by: Samplawski, Colin, et al.
Published: (2025)
Manifold-based Sampling for In-Context Hallucination Detection in Large Language Models
by: Vamshi, Bodla Krishna, et al.
Published: (2026)
by: Vamshi, Bodla Krishna, et al.
Published: (2026)
FCoReBench: Can Large Language Models Solve Challenging First-Order Combinatorial Reasoning Problems?
by: Mittal, Chinmay, et al.
Published: (2024)
by: Mittal, Chinmay, et al.
Published: (2024)
Certifying Knowledge Comprehension in LLMs
by: Chaudhary, Isha, et al.
Published: (2024)
by: Chaudhary, Isha, et al.
Published: (2024)
Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models
by: Lv, Ang, et al.
Published: (2024)
by: Lv, Ang, et al.
Published: (2024)
Thinking About Thinking: SAGE-nano's Inverse Reasoning for Self-Aware Language Models
by: Jha, Basab, et al.
Published: (2025)
by: Jha, Basab, et al.
Published: (2025)
Similar Items
-
Mem-Rec: Memory Efficient Recommendation System using Alternative Representation
by: Jha, Gopi Krishna, et al.
Published: (2023) -
TokenButler: Token Importance is Predictable
by: Akhauri, Yash, et al.
Published: (2025) -
Shears: Unstructured Sparsity with Neural Low-rank Adapter Search
by: Muñoz, J. Pablo, et al.
Published: (2024) -
SQFT: Low-cost Model Adaptation in Low-precision Sparse Foundation Models
by: Muñoz, Juan Pablo, et al.
Published: (2024) -
Low-Rank Adapters Meet Neural Architecture Search for LLM Compression
by: Muñoz, J. Pablo, et al.
Published: (2025)