LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM Inference
Fuente:
arXiv
Guardado en:
| Autores principales: | Mo, Zhiwen, Wang, Lei, Wei, Jianyu, Zeng, Zhichen, Cao, Shijie, Ma, Lingxiao, Jing, Naifeng, Cao, Ting, Xue, Jilong, Yang, Fan, Yang, Mao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
TENET: An Efficient Sparsity-Aware LUT-Centric Architecture for Ternary LLM Inference On Edge
por: Huang, Zhirui, et al.
Publicado: (2025)
por: Huang, Zhirui, et al.
Publicado: (2025)
LUT-DLA: Lookup Table as Efficient Extreme Low-Bit Deep Learning Accelerator
por: Li, Guoyu, et al.
Publicado: (2025)
por: Li, Guoyu, et al.
Publicado: (2025)
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
por: Du, Dayou, et al.
Publicado: (2025)
por: Du, Dayou, et al.
Publicado: (2025)
PolyLUT-Add: FPGA-based LUT Inference with Wide Inputs
por: Lou, Binglei, et al.
Publicado: (2024)
por: Lou, Binglei, et al.
Publicado: (2024)
T-MAN: Enabling End-to-End Low-Bit LLM Inference on NPUs via Unified Table Lookup
por: Wei, Jianyu, et al.
Publicado: (2025)
por: Wei, Jianyu, et al.
Publicado: (2025)
PolyLUT: Learning Piecewise Polynomials for Ultra-Low Latency FPGA LUT-based Inference
por: Andronic, Marta, et al.
Publicado: (2023)
por: Andronic, Marta, et al.
Publicado: (2023)
HGQ-LUT: Fast LUT-Aware Training and Efficient Architectures for DNN Inference
por: Sun, Chang, et al.
Publicado: (2026)
por: Sun, Chang, et al.
Publicado: (2026)
PolyLUT: Ultra-low Latency Polynomial Inference with Hardware-Aware Structured Pruning
por: Andronic, Marta, et al.
Publicado: (2025)
por: Andronic, Marta, et al.
Publicado: (2025)
Platinum: Path-Adaptable LUT-Based Accelerator Tailored for Low-Bit Weight Matrix Multiplication
por: Shan, Haoxuan, et al.
Publicado: (2025)
por: Shan, Haoxuan, et al.
Publicado: (2025)
LOCALUT: Harnessing Capacity-Computation Tradeoffs for LUT-Based Inference in DRAM-PIM
por: Hong, Junguk, et al.
Publicado: (2026)
por: Hong, Junguk, et al.
Publicado: (2026)
LUT-LLM: Efficient Large Language Model Inference with Memory-based Computations on FPGAs
por: He, Zifan, et al.
Publicado: (2025)
por: He, Zifan, et al.
Publicado: (2025)
DSLUT: An Asymmetric LUT and its Automatic Design Flow Based on Practical Functions
por: Yang, Moucheng, et al.
Publicado: (2025)
por: Yang, Moucheng, et al.
Publicado: (2025)
Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference
por: Hwang, Ranggi, et al.
Publicado: (2023)
por: Hwang, Ranggi, et al.
Publicado: (2023)
ReducedLUT: Table Decomposition with "Don't Care" Conditions
por: Cassidy, Oliver, et al.
Publicado: (2024)
por: Cassidy, Oliver, et al.
Publicado: (2024)
Enhancing LUT-based Deep Neural Networks Inference through Architecture and Connectivity Optimization
por: Lou, Binglei, et al.
Publicado: (2026)
por: Lou, Binglei, et al.
Publicado: (2026)
A Survey on LUT-based Deep Neural Networks Implemented in FPGAs
por: Guo, Zeyu
Publicado: (2025)
por: Guo, Zeyu
Publicado: (2025)
LUTMUL: Exceed Conventional FPGA Roofline Limit by LUT-based Efficient Multiplication for Neural Network Inference
por: Xie, Yanyue, et al.
Publicado: (2024)
por: Xie, Yanyue, et al.
Publicado: (2024)
Double Duty: FPGA Architecture to Enable Concurrent LUT and Adder Chain Usage
por: Pun, Junius, et al.
Publicado: (2025)
por: Pun, Junius, et al.
Publicado: (2025)
Ten-Four: An Open-Source Fused Dot Product Unit for Mixed-Precision GPGPU Tensor Cores
por: Rout, Nikhil, et al.
Publicado: (2025)
por: Rout, Nikhil, et al.
Publicado: (2025)
T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives
por: Pati, Suchita, et al.
Publicado: (2024)
por: Pati, Suchita, et al.
Publicado: (2024)
NeuraLUT: Hiding Neural Network Density in Boolean Synthesizable Functions
por: Andronic, Marta, et al.
Publicado: (2024)
por: Andronic, Marta, et al.
Publicado: (2024)
FlexiBit: Fully Flexible Precision Bit-parallel Accelerator Architecture for Arbitrary Mixed Precision AI
por: Tahmasebi, Faraz, et al.
Publicado: (2024)
por: Tahmasebi, Faraz, et al.
Publicado: (2024)
Architectural Isolation as a Timing Safety Primitive for Edge AI Medical Devices: Controlled Experimental Evidence on a Shared-Silicon Platform
por: Swami, Akul Mallayya
Publicado: (2026)
por: Swami, Akul Mallayya
Publicado: (2026)
TreeLUT: An Efficient Alternative to Deep Neural Networks for Inference Acceleration Using Gradient Boosted Decision Trees
por: Khataei, Alireza, et al.
Publicado: (2025)
por: Khataei, Alireza, et al.
Publicado: (2025)
CIMple: Standard-cell SRAM-based CIM with LUT-based split softmax for attention acceleration
por: Ahn, Bas, et al.
Publicado: (2026)
por: Ahn, Bas, et al.
Publicado: (2026)
Towards Efficient LUT-based PIM: A Scalable and Low-Power Approach for Modern Workloads
por: Khabbazan, Bahareh, et al.
Publicado: (2025)
por: Khabbazan, Bahareh, et al.
Publicado: (2025)
IPU: Flexible Hardware Introspection Units
por: McDougall, Ian, et al.
Publicado: (2023)
por: McDougall, Ian, et al.
Publicado: (2023)
SparseLUT: Sparse Connectivity Optimization for Lookup Table-based Deep Neural Networks
por: Lou, Binglei, et al.
Publicado: (2025)
por: Lou, Binglei, et al.
Publicado: (2025)
Hardware/Algorithm Co-design for Real-Time I/O Control with Improved Timing Accuracy and Robustness
por: Jiang, Zhe, et al.
Publicado: (2024)
por: Jiang, Zhe, et al.
Publicado: (2024)
MEDEA: A Design-Time Multi-Objective Manager for Energy-Efficient DNN Inference on Heterogeneous Ultra-Low Power Platforms
por: Taji, Hossein, et al.
Publicado: (2025)
por: Taji, Hossein, et al.
Publicado: (2025)
RISC-V Functional Safety for Autonomous Automotive Systems: An Analytical Framework and Research Roadmap for ML-Assisted Certification
por: Andreasyan, Nick, et al.
Publicado: (2026)
por: Andreasyan, Nick, et al.
Publicado: (2026)
OpenEye: A Scalable Open-Source Hardware Accelerator for DNNs
por: Lebold, Denis, et al.
Publicado: (2026)
por: Lebold, Denis, et al.
Publicado: (2026)
SAL-PIM: A Subarray-level Processing-in-Memory Architecture with LUT-based Linear Interpolation for Transformer-based Text Generation
por: Han, Wontak, et al.
Publicado: (2024)
por: Han, Wontak, et al.
Publicado: (2024)
KANELÉ: Kolmogorov-Arnold Networks for Efficient LUT-based Evaluation
por: Hoang, Duc, et al.
Publicado: (2025)
por: Hoang, Duc, et al.
Publicado: (2025)
Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM
por: Liu, Lian, et al.
Publicado: (2025)
por: Liu, Lian, et al.
Publicado: (2025)
GDEV-AI: A Generalized Evaluation of Deep Learning Inference Scaling and Architectural Saturation
por: Palaniappan, Kathiravan
Publicado: (2026)
por: Palaniappan, Kathiravan
Publicado: (2026)
Exploring GPU-to-GPU Communication: Insights into Supercomputer Interconnects
por: De Sensi, Daniele, et al.
Publicado: (2024)
por: De Sensi, Daniele, et al.
Publicado: (2024)
Towards Closing the Performance Gap for Cryptographic Kernels Between CPUs and Specialized Hardware
por: Zhang, Naifeng, et al.
Publicado: (2025)
por: Zhang, Naifeng, et al.
Publicado: (2025)
Fast NF4 Dequantization Kernels for Large Language Model Inference
por: Qi, Xiangbo, et al.
Publicado: (2026)
por: Qi, Xiangbo, et al.
Publicado: (2026)
A Survey on Hardware Accelerators for Large Language Models
por: Kachris, Christoforos
Publicado: (2024)
por: Kachris, Christoforos
Publicado: (2024)
Ejemplares similares
-
TENET: An Efficient Sparsity-Aware LUT-Centric Architecture for Ternary LLM Inference On Edge
por: Huang, Zhirui, et al.
Publicado: (2025) -
LUT-DLA: Lookup Table as Efficient Extreme Low-Bit Deep Learning Accelerator
por: Li, Guoyu, et al.
Publicado: (2025) -
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
por: Du, Dayou, et al.
Publicado: (2025) -
PolyLUT-Add: FPGA-based LUT Inference with Wide Inputs
por: Lou, Binglei, et al.
Publicado: (2024) -
T-MAN: Enabling End-to-End Low-Bit LLM Inference on NPUs via Unified Table Lookup
por: Wei, Jianyu, et al.
Publicado: (2025)