LUT-LLM: Efficient Large Language Model Inference with Memory-based Computations on FPGAs
Fuente:
arXiv
Saved in:
| Main Authors: | He, Zifan, Ye, Shengyu, Ma, Rui, Wang, Yang, Cong, Jason |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FlightLLM: Efficient Large Language Model Inference with a Complete Mapping Flow on FPGAs
by: Zeng, Shulin, et al.
Published: (2024)
by: Zeng, Shulin, et al.
Published: (2024)
LUT-DLA: Lookup Table as Efficient Extreme Low-Bit Deep Learning Accelerator
by: Li, Guoyu, et al.
Published: (2025)
by: Li, Guoyu, et al.
Published: (2025)
Leveraging Compute-in-Memory for Efficient Generative Model Inference in TPUs
by: Zhu, Zhantong, et al.
Published: (2025)
by: Zhu, Zhantong, et al.
Published: (2025)
PolyLUT-Add: FPGA-based LUT Inference with Wide Inputs
by: Lou, Binglei, et al.
Published: (2024)
by: Lou, Binglei, et al.
Published: (2024)
Enhancing LUT-based Deep Neural Networks Inference through Architecture and Connectivity Optimization
by: Lou, Binglei, et al.
Published: (2026)
by: Lou, Binglei, et al.
Published: (2026)
FlexLLM: Composable HLS Library for Flexible Hybrid LLM Accelerator Design
by: Zhang, Jiahao, et al.
Published: (2026)
by: Zhang, Jiahao, et al.
Published: (2026)
SWAT: Scalable and Efficient Window Attention-based Transformers Acceleration on FPGAs
by: Bai, Zhenyu, et al.
Published: (2024)
by: Bai, Zhenyu, et al.
Published: (2024)
MICSim: A Modular Simulator for Mixed-signal Compute-in-Memory based AI Accelerator
by: Wang, Cong, et al.
Published: (2024)
by: Wang, Cong, et al.
Published: (2024)
LUTMUL: Exceed Conventional FPGA Roofline Limit by LUT-based Efficient Multiplication for Neural Network Inference
by: Xie, Yanyue, et al.
Published: (2024)
by: Xie, Yanyue, et al.
Published: (2024)
TENET: An Efficient Sparsity-Aware LUT-Centric Architecture for Ternary LLM Inference On Edge
by: Huang, Zhirui, et al.
Published: (2025)
by: Huang, Zhirui, et al.
Published: (2025)
HLSTransform: Energy-Efficient Llama 2 Inference on FPGAs Via High Level Synthesis
by: He, Andy, et al.
Published: (2024)
by: He, Andy, et al.
Published: (2024)
SparseLUT: Sparse Connectivity Optimization for Lookup Table-based Deep Neural Networks
by: Lou, Binglei, et al.
Published: (2025)
by: Lou, Binglei, et al.
Published: (2025)
A Survey on LUT-based Deep Neural Networks Implemented in FPGAs
by: Guo, Zeyu
Published: (2025)
by: Guo, Zeyu
Published: (2025)
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
by: Fang, Yunhua, et al.
Published: (2025)
by: Fang, Yunhua, et al.
Published: (2025)
Efficient Deployment of CNN Models on Multiple In-Memory Computing Units
by: Bougioukou, Eleni, et al.
Published: (2025)
by: Bougioukou, Eleni, et al.
Published: (2025)
ELSA: An ELastic SNN Inference Architecture for Efficient Neuromorphic Computing
by: You, Kang, et al.
Published: (2026)
by: You, Kang, et al.
Published: (2026)
KVNAND: Efficient On-Device Large Language Model Inference Using DRAM-Free In-Flash Computing
by: Deng, Lishuo, et al.
Published: (2025)
by: Deng, Lishuo, et al.
Published: (2025)
LLM-DSE: Searching Accelerator Parameters with LLM Agents
by: Wang, Hanyu, et al.
Published: (2025)
by: Wang, Hanyu, et al.
Published: (2025)
AttentionLego: An Open-Source Building Block For Spatially-Scalable Large Language Model Accelerator With Processing-In-Memory Technology
by: Cong, Rongqing, et al.
Published: (2024)
by: Cong, Rongqing, et al.
Published: (2024)
Runtime Tunable Tsetlin Machines for Edge Inference on eFPGAs
by: Rahman, Tousif, et al.
Published: (2025)
by: Rahman, Tousif, et al.
Published: (2025)
Resource Utilization of Differentiable Logic Gate Networks Deployed on FPGAs
by: Wormald, Stephen, et al.
Published: (2026)
by: Wormald, Stephen, et al.
Published: (2026)
ROMA: a Read-Only-Memory-based Accelerator for QLoRA-based On-Device LLM
by: Wang, Wenqiang, et al.
Published: (2025)
by: Wang, Wenqiang, et al.
Published: (2025)
Pushing the Limits of BFP on Narrow Precision LLM Inference
by: Wang, Hui, et al.
Published: (2025)
by: Wang, Hui, et al.
Published: (2025)
LLM4SecHW: Leveraging Domain Specific Large Language Model for Hardware Debugging
by: Fu, Weimin, et al.
Published: (2024)
by: Fu, Weimin, et al.
Published: (2024)
HALO: Memory-Centric Heterogeneous Accelerator with 2.5D Integration for Low-Batch LLM Inference
by: Negi, Shubham, et al.
Published: (2025)
by: Negi, Shubham, et al.
Published: (2025)
Challenges and Research Directions for Large Language Model Inference Hardware
by: Ma, Xiaoyu, et al.
Published: (2026)
by: Ma, Xiaoyu, et al.
Published: (2026)
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference
by: Chen, Chun-Ting, et al.
Published: (2025)
by: Chen, Chun-Ting, et al.
Published: (2025)
Large Language Model (LLM) for Standard Cell Layout Design Optimization
by: Ho, Chia-Tung, et al.
Published: (2024)
by: Ho, Chia-Tung, et al.
Published: (2024)
Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference
by: Wang, Xinyu, et al.
Published: (2026)
by: Wang, Xinyu, et al.
Published: (2026)
Efficient Calibration for RRAM-based In-Memory Computing using DoRA
by: Dong, Weirong, et al.
Published: (2025)
by: Dong, Weirong, et al.
Published: (2025)
NeuralMatrix: Compute the Entire Neural Networks with Linear Matrix Operations for Efficient Inference
by: Sun, Ruiqi, et al.
Published: (2023)
by: Sun, Ruiqi, et al.
Published: (2023)
PrefixAgent: An LLM-Powered Design Framework for Efficient Prefix Adder Optimization
by: Zuo, Dongsheng, et al.
Published: (2025)
by: Zuo, Dongsheng, et al.
Published: (2025)
FractalCloud: A Fractal-Inspired Architecture for Efficient Large-Scale Point Cloud Processing
by: Fu, Yuzhe, et al.
Published: (2025)
by: Fu, Yuzhe, et al.
Published: (2025)
LLM4EDA: Emerging Progress in Large Language Models for Electronic Design Automation
by: Zhong, Ruizhe, et al.
Published: (2023)
by: Zhong, Ruizhe, et al.
Published: (2023)
PIMphony: Overcoming Bandwidth and Capacity Inefficiency in PIM-based Long-Context LLM Inference System
by: Kwon, Hyucksung, et al.
Published: (2024)
by: Kwon, Hyucksung, et al.
Published: (2024)
EvoVerilog: Large Langugage Model Assisted Evolution of Verilog Code
by: Guo, Ping, et al.
Published: (2025)
by: Guo, Ping, et al.
Published: (2025)
Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache Management Beyond GPU Limits
by: Kim, Dowon, et al.
Published: (2025)
by: Kim, Dowon, et al.
Published: (2025)
HPIM: Heterogeneous Processing-In-Memory-based Accelerator for Large Language Models Inference
by: Duan, Cenlin, et al.
Published: (2025)
by: Duan, Cenlin, et al.
Published: (2025)
YOCO: A Hybrid In-Memory Computing Architecture with 8-bit Sub-PetaOps/W In-Situ Multiply Arithmetic for Large-Scale AI
by: Xuan, Zihao, et al.
Published: (2023)
by: Xuan, Zihao, et al.
Published: (2023)
FAMOUS: Flexible Accelerator for the Attention Mechanism of Transformer on UltraScale+ FPGAs
by: Kabir, Ehsan, et al.
Published: (2024)
by: Kabir, Ehsan, et al.
Published: (2024)
Similar Items
-
FlightLLM: Efficient Large Language Model Inference with a Complete Mapping Flow on FPGAs
by: Zeng, Shulin, et al.
Published: (2024) -
LUT-DLA: Lookup Table as Efficient Extreme Low-Bit Deep Learning Accelerator
by: Li, Guoyu, et al.
Published: (2025) -
Leveraging Compute-in-Memory for Efficient Generative Model Inference in TPUs
by: Zhu, Zhantong, et al.
Published: (2025) -
PolyLUT-Add: FPGA-based LUT Inference with Wide Inputs
by: Lou, Binglei, et al.
Published: (2024) -
Enhancing LUT-based Deep Neural Networks Inference through Architecture and Connectivity Optimization
by: Lou, Binglei, et al.
Published: (2026)