T-SAR: A Full-Stack Co-design for CPU-Only Ternary LLM Inference via In-Place SIMD ALU Reorganization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Oh, Hyunwoo, Nam, KyungIn, Bhattacharjya, Rajat, Chen, Hanning, Das, Tamoghno, Yun, Sanggeon, Jang, Suyeon, Ding, Andrew, Dutt, Nikil, Imani, Mohsen
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908659723796480
author Oh, Hyunwoo
Nam, KyungIn
Bhattacharjya, Rajat
Chen, Hanning
Das, Tamoghno
Yun, Sanggeon
Jang, Suyeon
Ding, Andrew
Dutt, Nikil
Imani, Mohsen
author_facet Oh, Hyunwoo
Nam, KyungIn
Bhattacharjya, Rajat
Chen, Hanning
Das, Tamoghno
Yun, Sanggeon
Jang, Suyeon
Ding, Andrew
Dutt, Nikil
Imani, Mohsen
contents Recent advances in LLMs have outpaced the computational and memory capacities of edge platforms that primarily employ CPUs, thereby challenging efficient and scalable deployment. While ternary quantization enables significant resource savings, existing CPU solutions rely heavily on memory-based lookup tables (LUTs) which limit scalability, and FPGA or GPU accelerators remain impractical for edge use. This paper presents T-SAR, the first framework to achieve scalable ternary LLM inference on CPUs by repurposing the SIMD register file for dynamic, in-register LUT generation with minimal hardware modifications. T-SAR eliminates memory bottlenecks and maximizes data-level parallelism, delivering 5.6-24.5x and 1.1-86.2x improvements in GEMM latency and GEMV throughput, respectively, with only 3.2% power and 1.4% area overheads in SIMD units. T-SAR achieves up to 2.5-4.9x the energy efficiency of an NVIDIA Jetson AGX Orin, establishing a practical approach for efficient LLM inference on edge platforms.
format Preprint
id arxiv_https___arxiv_org_abs_2511_13676
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle T-SAR: A Full-Stack Co-design for CPU-Only Ternary LLM Inference via In-Place SIMD ALU Reorganization
Oh, Hyunwoo
Nam, KyungIn
Bhattacharjya, Rajat
Chen, Hanning
Das, Tamoghno
Yun, Sanggeon
Jang, Suyeon
Ding, Andrew
Dutt, Nikil
Imani, Mohsen
Hardware Architecture
Machine Learning
Recent advances in LLMs have outpaced the computational and memory capacities of edge platforms that primarily employ CPUs, thereby challenging efficient and scalable deployment. While ternary quantization enables significant resource savings, existing CPU solutions rely heavily on memory-based lookup tables (LUTs) which limit scalability, and FPGA or GPU accelerators remain impractical for edge use. This paper presents T-SAR, the first framework to achieve scalable ternary LLM inference on CPUs by repurposing the SIMD register file for dynamic, in-register LUT generation with minimal hardware modifications. T-SAR eliminates memory bottlenecks and maximizes data-level parallelism, delivering 5.6-24.5x and 1.1-86.2x improvements in GEMM latency and GEMV throughput, respectively, with only 3.2% power and 1.4% area overheads in SIMD units. T-SAR achieves up to 2.5-4.9x the energy efficiency of an NVIDIA Jetson AGX Orin, establishing a practical approach for efficient LLM inference on edge platforms.
title T-SAR: A Full-Stack Co-design for CPU-Only Ternary LLM Inference via In-Place SIMD ALU Reorganization
topic Hardware Architecture
Machine Learning
url https://arxiv.org/abs/2511.13676