GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kang, Hao, Zhang, Qingru, Kundu, Souvik, Jeong, Geonhwa, Liu, Zaoxing, Krishna, Tushar, Zhao, Tuo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!