Residual vector quantization for KV cache compression in large language model
Fuente:
arXiv
Guardado en:
| Autor principal: | Kumar, Ankur |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
QET: Enhancing Quantized LLM Parameters and KV cache Compression through Element Substitution and Residual Clustering
por: Wang, Yanshu, et al.
Publicado: (2024)
por: Wang, Yanshu, et al.
Publicado: (2024)
An experimental study of KV cache reuse strategies in chunk-level caching systems
por: Cestola, Samuel, et al.
Publicado: (2026)
por: Cestola, Samuel, et al.
Publicado: (2026)
SALS: Sparse Attention in Latent Space for KV cache Compression
por: Mu, Junlin, et al.
Publicado: (2025)
por: Mu, Junlin, et al.
Publicado: (2025)
Layer-wise dynamic rank for compressing large language models
por: Mi, Zhendong, et al.
Publicado: (2025)
por: Mi, Zhendong, et al.
Publicado: (2025)
NIRVANA: Structured pruning reimagined for large language models compression
por: Ai, Mengting, et al.
Publicado: (2025)
por: Ai, Mengting, et al.
Publicado: (2025)
BaKlaVa -- Budgeted Allocation of KV cache for Long-context Inference
por: Gulhan, Ahmed Burak, et al.
Publicado: (2025)
por: Gulhan, Ahmed Burak, et al.
Publicado: (2025)
IceCache: Memory-efficient KV-cache Management for Long-Sequence LLMs
por: Mao, Yuzhen, et al.
Publicado: (2026)
por: Mao, Yuzhen, et al.
Publicado: (2026)
KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent Systems
por: Ye, Hancheng, et al.
Publicado: (2025)
por: Ye, Hancheng, et al.
Publicado: (2025)
A general tensor-structured compression scheme for efficient large language models
por: Lu, Ying, et al.
Publicado: (2026)
por: Lu, Ying, et al.
Publicado: (2026)
Query-efficient model evaluation using cached responses
por: Helm, Hayden, et al.
Publicado: (2026)
por: Helm, Hayden, et al.
Publicado: (2026)
Residual-Mass Accounting for Partial-KV Decoding
por: Hoshi, Yasuto, et al.
Publicado: (2026)
por: Hoshi, Yasuto, et al.
Publicado: (2026)
Individualized non-uniform quantization for vector search
por: Tepper, Mariano, et al.
Publicado: (2025)
por: Tepper, Mariano, et al.
Publicado: (2025)
Visual cognition in multimodal large language models
por: Buschoff, Luca M. Schulze, et al.
Publicado: (2023)
por: Buschoff, Luca M. Schulze, et al.
Publicado: (2023)
Hypothesis generation and updating in large language models
por: Xiong, Hua-Dong
Publicado: (2026)
por: Xiong, Hua-Dong
Publicado: (2026)
Representation in large language models
por: Yetman, Cameron
Publicado: (2025)
por: Yetman, Cameron
Publicado: (2025)
Price of universality in vector quantization is at most 0.11 bit
por: Harbuzova, Alina, et al.
Publicado: (2026)
por: Harbuzova, Alina, et al.
Publicado: (2026)
Amortizing intractable inference in large language models
por: Hu, Edward J., et al.
Publicado: (2023)
por: Hu, Edward J., et al.
Publicado: (2023)
Training microrobots to swim by a large language model
por: Xu, Zhuoqun, et al.
Publicado: (2024)
por: Xu, Zhuoqun, et al.
Publicado: (2024)
OCTOPUS: Optimized KV Cache for Transformers via Octahedral Parametrization Under optimal Squared error quantization
por: Boss, Mark, et al.
Publicado: (2026)
por: Boss, Mark, et al.
Publicado: (2026)
Alignment faking in large language models
por: Greenblatt, Ryan, et al.
Publicado: (2024)
por: Greenblatt, Ryan, et al.
Publicado: (2024)
Latent Phase-Shift Rollback: Inference-Time Error Correction via Residual Stream Monitoring and KV-Cache Steering
por: Gupta, Manan, et al.
Publicado: (2026)
por: Gupta, Manan, et al.
Publicado: (2026)
AI-AI Bias: large language models favor communications generated by large language models
por: Laurito, Walter, et al.
Publicado: (2024)
por: Laurito, Walter, et al.
Publicado: (2024)
Variational quantization for state space models
por: David, Etienne, et al.
Publicado: (2024)
por: David, Etienne, et al.
Publicado: (2024)
The Residual Stream Is All You Need: On the Redundancy of the KV Cache in Transformer Inference
por: Qasim, Kaleem Ullah, et al.
Publicado: (2026)
por: Qasim, Kaleem Ullah, et al.
Publicado: (2026)
Quantifying construct validity in large language model evaluations
por: Kearns, Ryan Othniel
Publicado: (2026)
por: Kearns, Ryan Othniel
Publicado: (2026)
Long-form factuality in large language models
por: Wei, Jerry, et al.
Publicado: (2024)
por: Wei, Jerry, et al.
Publicado: (2024)
Can large language models explore in-context?
por: Krishnamurthy, Akshay, et al.
Publicado: (2024)
por: Krishnamurthy, Akshay, et al.
Publicado: (2024)
Quantifying perturbation impacts for large language models
por: Rauba, Paulius, et al.
Publicado: (2024)
por: Rauba, Paulius, et al.
Publicado: (2024)
Alignment of large language models with constrained learning
por: Zhang, Botong, et al.
Publicado: (2025)
por: Zhang, Botong, et al.
Publicado: (2025)
Insights into a radiology-specialised multimodal large language model with sparse autoencoders
por: Bouzid, Kenza, et al.
Publicado: (2025)
por: Bouzid, Kenza, et al.
Publicado: (2025)
ARETE: an R package for Automated REtrieval from TExt with large language models
por: Branco, Vasco V., et al.
Publicado: (2025)
por: Branco, Vasco V., et al.
Publicado: (2025)
Uniform error bounds for quantized dynamical models
por: Metakalard, Abdelkader, et al.
Publicado: (2026)
por: Metakalard, Abdelkader, et al.
Publicado: (2026)
Less can be more for predicting properties with large language models
por: Alampara, Nawaf, et al.
Publicado: (2024)
por: Alampara, Nawaf, et al.
Publicado: (2024)
Are large language models superhuman chemists?
por: Mirza, Adrian, et al.
Publicado: (2024)
por: Mirza, Adrian, et al.
Publicado: (2024)
Harnessing large-language models to generate private synthetic text
por: Kurakin, Alexey, et al.
Publicado: (2023)
por: Kurakin, Alexey, et al.
Publicado: (2023)
Prompt reinforcing for long-term planning of large language models
por: Lin, Hsien-Chin, et al.
Publicado: (2025)
por: Lin, Hsien-Chin, et al.
Publicado: (2025)
Relational reasoning and inductive bias in transformers and large language models
por: Geerts, Jesse, et al.
Publicado: (2025)
por: Geerts, Jesse, et al.
Publicado: (2025)
TOKON: TOKenization-Optimized Normalization for time series analysis with a large language model
por: Yang, Janghoon
Publicado: (2025)
por: Yang, Janghoon
Publicado: (2025)
AngelSlim: A more accessible, comprehensive, and efficient toolkit for large model compression
por: Cen, Rui, et al.
Publicado: (2026)
por: Cen, Rui, et al.
Publicado: (2026)
Efficiency optimization of large-scale language models based on deep learning in natural language processing tasks
por: Mei, Taiyuan, et al.
Publicado: (2024)
por: Mei, Taiyuan, et al.
Publicado: (2024)
Ejemplares similares
-
QET: Enhancing Quantized LLM Parameters and KV cache Compression through Element Substitution and Residual Clustering
por: Wang, Yanshu, et al.
Publicado: (2024) -
An experimental study of KV cache reuse strategies in chunk-level caching systems
por: Cestola, Samuel, et al.
Publicado: (2026) -
SALS: Sparse Attention in Latent Space for KV cache Compression
por: Mu, Junlin, et al.
Publicado: (2025) -
Layer-wise dynamic rank for compressing large language models
por: Mi, Zhendong, et al.
Publicado: (2025) -
NIRVANA: Structured pruning reimagined for large language models compression
por: Ai, Mengting, et al.
Publicado: (2025)