LUTMUL: Exceed Conventional FPGA Roofline Limit by LUT-based Efficient Multiplication for Neural Network Inference
Fuente:
arXiv
Guardado en:
| Autores principales: | Xie, Yanyue, Li, Zhengang, Diaconu, Dana, Handagala, Suranga, Leeser, Miriam, Lin, Xue |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PolyLUT: Learning Piecewise Polynomials for Ultra-Low Latency FPGA LUT-based Inference
por: Andronic, Marta, et al.
Publicado: (2023)
por: Andronic, Marta, et al.
Publicado: (2023)
PolyLUT-Add: FPGA-based LUT Inference with Wide Inputs
por: Lou, Binglei, et al.
Publicado: (2024)
por: Lou, Binglei, et al.
Publicado: (2024)
Memory-efficient Sketch Acceleration for Handling Large Network Flows on FPGAs
por: Han, Zhaoyang, et al.
Publicado: (2025)
por: Han, Zhaoyang, et al.
Publicado: (2025)
Extracting TCPIP Headers at High Speed for the Anonymized Network Traffic Graph Challenge
por: Han, Zhaoyang, et al.
Publicado: (2024)
por: Han, Zhaoyang, et al.
Publicado: (2024)
Double Duty: FPGA Architecture to Enable Concurrent LUT and Adder Chain Usage
por: Pun, Junius, et al.
Publicado: (2025)
por: Pun, Junius, et al.
Publicado: (2025)
TENET: An Efficient Sparsity-Aware LUT-Centric Architecture for Ternary LLM Inference On Edge
por: Huang, Zhirui, et al.
Publicado: (2025)
por: Huang, Zhirui, et al.
Publicado: (2025)
A Survey on LUT-based Deep Neural Networks Implemented in FPGAs
por: Guo, Zeyu
Publicado: (2025)
por: Guo, Zeyu
Publicado: (2025)
TreeLUT: An Efficient Alternative to Deep Neural Networks for Inference Acceleration Using Gradient Boosted Decision Trees
por: Khataei, Alireza, et al.
Publicado: (2025)
por: Khataei, Alireza, et al.
Publicado: (2025)
Enhancing LUT-based Deep Neural Networks Inference through Architecture and Connectivity Optimization
por: Lou, Binglei, et al.
Publicado: (2026)
por: Lou, Binglei, et al.
Publicado: (2026)
Energy-Efficient FPGA Framework for Non-Quantized Convolutional Neural Networks
por: Athanasiadis, Angelos, et al.
Publicado: (2025)
por: Athanasiadis, Angelos, et al.
Publicado: (2025)
LOCALUT: Harnessing Capacity-Computation Tradeoffs for LUT-Based Inference in DRAM-PIM
por: Hong, Junguk, et al.
Publicado: (2026)
por: Hong, Junguk, et al.
Publicado: (2026)
TerEffic: Highly Efficient Ternary LLM Inference on FPGA
por: Yin, Chenyang, et al.
Publicado: (2025)
por: Yin, Chenyang, et al.
Publicado: (2025)
TROOP: At-the-Roofline Performance for Vector Processors on Low Operational Intensity Workloads
por: Purayil, Navaneeth Kunhi, et al.
Publicado: (2025)
por: Purayil, Navaneeth Kunhi, et al.
Publicado: (2025)
Platinum: Path-Adaptable LUT-Based Accelerator Tailored for Low-Bit Weight Matrix Multiplication
por: Shan, Haoxuan, et al.
Publicado: (2025)
por: Shan, Haoxuan, et al.
Publicado: (2025)
SpecMamba: Accelerating Mamba Inference on FPGA with Speculative Decoding
por: Zhong, Linfeng, et al.
Publicado: (2025)
por: Zhong, Linfeng, et al.
Publicado: (2025)
LUT-LLM: Efficient Large Language Model Inference with Memory-based Computations on FPGAs
por: He, Zifan, et al.
Publicado: (2025)
por: He, Zifan, et al.
Publicado: (2025)
NeuraLUT: Hiding Neural Network Density in Boolean Synthesizable Functions
por: Andronic, Marta, et al.
Publicado: (2024)
por: Andronic, Marta, et al.
Publicado: (2024)
XtraMAC: An Efficient MAC Architecture for Mixed-Precision LLM Inference on FPGA
por: Yu, Feng, et al.
Publicado: (2026)
por: Yu, Feng, et al.
Publicado: (2026)
Binary Neural Network Implementation for Handwritten Digit Recognition on FPGA
por: Ertörer, Emir Devlet, et al.
Publicado: (2025)
por: Ertörer, Emir Devlet, et al.
Publicado: (2025)
HGQ-LUT: Fast LUT-Aware Training and Efficient Architectures for DNN Inference
por: Sun, Chang, et al.
Publicado: (2026)
por: Sun, Chang, et al.
Publicado: (2026)
Towards Efficient LUT-based PIM: A Scalable and Low-Power Approach for Modern Workloads
por: Khabbazan, Bahareh, et al.
Publicado: (2025)
por: Khabbazan, Bahareh, et al.
Publicado: (2025)
Pushing up to the Limit of Memory Bandwidth and Capacity Utilization for Efficient LLM Decoding on Embedded FPGA
por: Li, Jindong, et al.
Publicado: (2025)
por: Li, Jindong, et al.
Publicado: (2025)
An Efficient Hardware Implementation of Elliptic Curve Point Multiplication over $GF(2^m)$ on FPGA
por: Kumari, Ruby, et al.
Publicado: (2025)
por: Kumari, Ruby, et al.
Publicado: (2025)
EMiX: Emulating Beyond Single-FPGA Limits
por: Kropotov, Alexander, et al.
Publicado: (2026)
por: Kropotov, Alexander, et al.
Publicado: (2026)
SparseLUT: Sparse Connectivity Optimization for Lookup Table-based Deep Neural Networks
por: Lou, Binglei, et al.
Publicado: (2025)
por: Lou, Binglei, et al.
Publicado: (2025)
DECA: A Near-Core LLM Decompression Accelerator Grounded on a 3D Roofline Model
por: Gerogiannis, Gerasimos, et al.
Publicado: (2025)
por: Gerogiannis, Gerasimos, et al.
Publicado: (2025)
Embedded FPGA Acceleration of Brain-Like Neural Networks: Online Learning to Scalable Inference
por: Hafiz, Muhammad Ihsan Al, et al.
Publicado: (2025)
por: Hafiz, Muhammad Ihsan Al, et al.
Publicado: (2025)
Table-Lookup MAC: Scalable Processing of Quantised Neural Networks in FPGA Soft Logic
por: Gerlinghoff, Daniel, et al.
Publicado: (2024)
por: Gerlinghoff, Daniel, et al.
Publicado: (2024)
SIRA: Scaled-Integer Range Analysis for Optimizing FPGA Dataflow Neural Network Accelerators
por: Umuroglu, Yaman, et al.
Publicado: (2025)
por: Umuroglu, Yaman, et al.
Publicado: (2025)
SuperFlow: A Fully-Customized RTL-to-GDS Design Automation Flow for Adiabatic Quantum-Flux-Parametron Superconducting Circuits
por: Xie, Yanyue, et al.
Publicado: (2024)
por: Xie, Yanyue, et al.
Publicado: (2024)
Systolic Array Data Flows for Efficient Matrix Multiplication in Deep Neural Networks
por: Raja, Tejas
Publicado: (2024)
por: Raja, Tejas
Publicado: (2024)
Real-Time Adaptive Neural Network on FPGA: Enhancing Adaptability through Dynamic Classifier Selection
por: Bouazzaoui, Achraf El, et al.
Publicado: (2023)
por: Bouazzaoui, Achraf El, et al.
Publicado: (2023)
FireFly-P: FPGA-Accelerated Spiking Neural Network Plasticity for Robust Adaptive Control
por: Li, Tenglong, et al.
Publicado: (2026)
por: Li, Tenglong, et al.
Publicado: (2026)
PolyLUT: Ultra-low Latency Polynomial Inference with Hardware-Aware Structured Pruning
por: Andronic, Marta, et al.
Publicado: (2025)
por: Andronic, Marta, et al.
Publicado: (2025)
DSLUT: An Asymmetric LUT and its Automatic Design Flow Based on Practical Functions
por: Yang, Moucheng, et al.
Publicado: (2025)
por: Yang, Moucheng, et al.
Publicado: (2025)
Evaluating Four FPGA-accelerated Space Use Cases based on Neural Network Algorithms for On-board Inference
por: Antunes, Pedro, et al.
Publicado: (2026)
por: Antunes, Pedro, et al.
Publicado: (2026)
RooflineBench: A Benchmarking Framework for On-Device LLMs via Roofline Analysis
por: Bi, Zhen, et al.
Publicado: (2026)
por: Bi, Zhen, et al.
Publicado: (2026)
Efficient and Accurate Graph Classification with Hyperdimensional Computing on FPGA
por: Arockiaraj, Jebacyril, et al.
Publicado: (2025)
por: Arockiaraj, Jebacyril, et al.
Publicado: (2025)
Online Training and Inference System on Edge FPGA Using Delayed Feedback Reservoir
por: Ikeda, Sosei, et al.
Publicado: (2025)
por: Ikeda, Sosei, et al.
Publicado: (2025)
SpeedLLM: An FPGA Co-design of Large Language Model Inference Accelerator
por: Wang, Peipei, et al.
Publicado: (2025)
por: Wang, Peipei, et al.
Publicado: (2025)
Ejemplares similares
-
PolyLUT: Learning Piecewise Polynomials for Ultra-Low Latency FPGA LUT-based Inference
por: Andronic, Marta, et al.
Publicado: (2023) -
PolyLUT-Add: FPGA-based LUT Inference with Wide Inputs
por: Lou, Binglei, et al.
Publicado: (2024) -
Memory-efficient Sketch Acceleration for Handling Large Network Flows on FPGAs
por: Han, Zhaoyang, et al.
Publicado: (2025) -
Extracting TCPIP Headers at High Speed for the Anonymized Network Traffic Graph Challenge
por: Han, Zhaoyang, et al.
Publicado: (2024) -
Double Duty: FPGA Architecture to Enable Concurrent LUT and Adder Chain Usage
por: Pun, Junius, et al.
Publicado: (2025)