EVA: Recasting LLM Decoding into GEMM via an Efficient Vector Quantization Architecture
Fuente:
Zenodo
Enregistré dans:
| Auteurs principaux: | , , , , , , , , , , |
|---|---|
| Format: | Recurso digital |
| Publié: |
Zenodo
2026
|
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866902182331154432 |
|---|---|
| author | Duan, Bowen Guo, Cong Wei, Chiyue Shan, Haoxuan Fu, Yuzhe Chen, Xinhua Xu, Yifan Zhang, Ziyue Zhou, Changchun Li, Hai Chen, Yiran |
| author_facet | Duan, Bowen Guo, Cong Wei, Chiyue Shan, Haoxuan Fu, Yuzhe Chen, Xinhua Xu, Yifan Zhang, Ziyue Zhou, Changchun Li, Hai Chen, Yiran |
| contents | <div> <div dir="ltr"> <p>This repository provides the official implementation and artifacts for the ISCA 2026 paper "EVA: Recasting LLM Decoding into GEMM via an Efficient Vector Quantization Architecture."</p> <p>This release corresponds to the artifact-evaluated version of the codebase. It includes all scripts, configuration files, and Jupyter notebooks required to reproduce the hardware performance and algorithm accuracy results reported in the paper.</p> </div> </div> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_19444241 |
| institution | Zenodo |
| language | |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | EVA: Recasting LLM Decoding into GEMM via an Efficient Vector Quantization Architecture Duan, Bowen Guo, Cong Wei, Chiyue Shan, Haoxuan Fu, Yuzhe Chen, Xinhua Xu, Yifan Zhang, Ziyue Zhou, Changchun Li, Hai Chen, Yiran <div> <div dir="ltr"> <p>This repository provides the official implementation and artifacts for the ISCA 2026 paper "EVA: Recasting LLM Decoding into GEMM via an Efficient Vector Quantization Architecture."</p> <p>This release corresponds to the artifact-evaluated version of the codebase. It includes all scripts, configuration files, and Jupyter notebooks required to reproduce the hardware performance and algorithm accuracy results reported in the paper.</p> </div> </div> |
| title | EVA: Recasting LLM Decoding into GEMM via an Efficient Vector Quantization Architecture |
| url | https://doi.org/10.5281/zenodo.19444241 |