Fast and Efficient 2-bit LLM Inference on GPU: 2/4/16-bit in a Weight Matrix with Asynchronous Dequantization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Jinhao, Xu, Jiaming, Li, Shiyao, Huang, Shan, Liu, Jun, Lian, Yaoxiu, Dai, Guohao
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!

Similar Items