EVA: Recasting LLM Decoding into GEMM via an Efficient Vector Quantization Architecture

Fuente: Zenodo
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Duan, Bowen, Guo, Cong, Wei, Chiyue, Shan, Haoxuan, Fu, Yuzhe, Chen, Xinhua, Xu, Yifan, Zhang, Ziyue, Zhou, Changchun, Li, Hai, Chen, Yiran
Format: Recurso digital
Publié: Zenodo 2026
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866902182331154432
author Duan, Bowen
Guo, Cong
Wei, Chiyue
Shan, Haoxuan
Fu, Yuzhe
Chen, Xinhua
Xu, Yifan
Zhang, Ziyue
Zhou, Changchun
Li, Hai
Chen, Yiran
author_facet Duan, Bowen
Guo, Cong
Wei, Chiyue
Shan, Haoxuan
Fu, Yuzhe
Chen, Xinhua
Xu, Yifan
Zhang, Ziyue
Zhou, Changchun
Li, Hai
Chen, Yiran
contents <div> <div dir="ltr"> <p>This repository provides the official implementation and artifacts for the ISCA 2026 paper "EVA: Recasting LLM Decoding into GEMM via an Efficient Vector Quantization Architecture."</p> <p>This release corresponds to the artifact-evaluated version of the codebase. It includes all scripts, configuration files, and Jupyter notebooks required to reproduce the hardware performance and algorithm accuracy results reported in the paper.</p> </div> </div>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_19444241
institution Zenodo
language
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle EVA: Recasting LLM Decoding into GEMM via an Efficient Vector Quantization Architecture
Duan, Bowen
Guo, Cong
Wei, Chiyue
Shan, Haoxuan
Fu, Yuzhe
Chen, Xinhua
Xu, Yifan
Zhang, Ziyue
Zhou, Changchun
Li, Hai
Chen, Yiran
<div> <div dir="ltr"> <p>This repository provides the official implementation and artifacts for the ISCA 2026 paper "EVA: Recasting LLM Decoding into GEMM via an Efficient Vector Quantization Architecture."</p> <p>This release corresponds to the artifact-evaluated version of the codebase. It includes all scripts, configuration files, and Jupyter notebooks required to reproduce the hardware performance and algorithm accuracy results reported in the paper.</p> </div> </div>
title EVA: Recasting LLM Decoding into GEMM via an Efficient Vector Quantization Architecture
url https://doi.org/10.5281/zenodo.19444241