KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Tianyi, Yi, Jonah, Xu, Zhaozhuo, Shrivastava, Anshumali
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!

Similar Items