PolyLUT: Ultra-low Latency Polynomial Inference with Hardware-Aware Structured Pruning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Andronic, Marta, Li, Jiawen, Constantinides, George A. |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
PolyLUT: Learning Piecewise Polynomials for Ultra-Low Latency FPGA LUT-based Inference
par: Andronic, Marta, et autres
Publié: (2023)
par: Andronic, Marta, et autres
Publié: (2023)
NeuraLUT: Hiding Neural Network Density in Boolean Synthesizable Functions
par: Andronic, Marta, et autres
Publié: (2024)
par: Andronic, Marta, et autres
Publié: (2024)
PolyLUT-Add: FPGA-based LUT Inference with Wide Inputs
par: Lou, Binglei, et autres
Publié: (2024)
par: Lou, Binglei, et autres
Publié: (2024)
ReducedLUT: Table Decomposition with "Don't Care" Conditions
par: Cassidy, Oliver, et autres
Publié: (2024)
par: Cassidy, Oliver, et autres
Publié: (2024)
ATHEENA: A Toolflow for Hardware Early-Exit Network Automation
par: Biggs, Benjamin, et autres
Publié: (2023)
par: Biggs, Benjamin, et autres
Publié: (2023)
BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration
par: Chen, Yuzong, et autres
Publié: (2024)
par: Chen, Yuzong, et autres
Publié: (2024)
NeuraLUT-Assemble: Hardware-aware Assembling of Sub-Neural Networks for Efficient LUT Inference
par: Andronic, Marta, et autres
Publié: (2025)
par: Andronic, Marta, et autres
Publié: (2025)
HGQ-LUT: Fast LUT-Aware Training and Efficient Architectures for DNN Inference
par: Sun, Chang, et autres
Publié: (2026)
par: Sun, Chang, et autres
Publié: (2026)
Low Power Vision Transformer Accelerator with Hardware-Aware Pruning and Optimized Dataflow
par: Hsiung, Ching-Lin, et autres
Publié: (2025)
par: Hsiung, Ching-Lin, et autres
Publié: (2025)
AMPLE: Event-Driven Accelerator for Mixed-Precision Inference of Graph Neural Networks
par: Gimenes, Pedro, et autres
Publié: (2025)
par: Gimenes, Pedro, et autres
Publié: (2025)
TreeLUT: An Efficient Alternative to Deep Neural Networks for Inference Acceleration Using Gradient Boosted Decision Trees
par: Khataei, Alireza, et autres
Publié: (2025)
par: Khataei, Alireza, et autres
Publié: (2025)
Exploring FPGA designs for MX and beyond
par: Samson, Ebby, et autres
Publié: (2024)
par: Samson, Ebby, et autres
Publié: (2024)
SigmaQuant: Hardware-Aware Heterogeneous Quantization Method for Edge DNN Inference
par: Liu, Qunyou, et autres
Publié: (2026)
par: Liu, Qunyou, et autres
Publié: (2026)
Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective
par: Li, Jinhao, et autres
Publié: (2024)
par: Li, Jinhao, et autres
Publié: (2024)
LUTMUL: Exceed Conventional FPGA Roofline Limit by LUT-based Efficient Multiplication for Neural Network Inference
par: Xie, Yanyue, et autres
Publié: (2024)
par: Xie, Yanyue, et autres
Publié: (2024)
ALADIN: Accuracy-Latency-Aware Design-space Inference Analysis for Embedded AI Accelerators
par: Baldi, T., et autres
Publié: (2026)
par: Baldi, T., et autres
Publié: (2026)
Energy-Aware Deep Learning on Resource-Constrained Hardware
par: Millar, Josh, et autres
Publié: (2025)
par: Millar, Josh, et autres
Publié: (2025)
FPGA Resource-aware Structured Pruning for Real-Time Neural Networks
par: Ramhorst, Benjamin, et autres
Publié: (2023)
par: Ramhorst, Benjamin, et autres
Publié: (2023)
PolyThrottle: Energy-efficient Neural Network Inference on Edge Devices
par: Yan, Minghao, et autres
Publié: (2023)
par: Yan, Minghao, et autres
Publié: (2023)
LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM Inference
par: Mo, Zhiwen, et autres
Publié: (2024)
par: Mo, Zhiwen, et autres
Publié: (2024)
Architectural Implications of Neural Network Inference for High Data-Rate, Low-Latency Scientific Applications
par: Weng, Olivia, et autres
Publié: (2024)
par: Weng, Olivia, et autres
Publié: (2024)
HASS: Hardware-Aware Sparsity Search for Dataflow DNN Accelerator
par: Yu, Zhewen, et autres
Publié: (2024)
par: Yu, Zhewen, et autres
Publié: (2024)
CATransformers: Carbon Aware Transformers Through Joint Model-Hardware Optimization
par: Wang, Irene, et autres
Publié: (2025)
par: Wang, Irene, et autres
Publié: (2025)
Hardware-Aware Neural Dropout Search for Reliable Uncertainty Prediction on FPGA
par: Zhang, Zehuan, et autres
Publié: (2024)
par: Zhang, Zehuan, et autres
Publié: (2024)
On-Device Qwen2.5: Efficient LLM Inference with Model Compression and Hardware Acceleration
par: Xiang, Maoyang, et autres
Publié: (2025)
par: Xiang, Maoyang, et autres
Publié: (2025)
SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
par: Wang, Wenxun, et autres
Publié: (2025)
par: Wang, Wenxun, et autres
Publié: (2025)
Active Imitation Learning for Thermal- and Kernel-Aware LFM Inference on 3D S-NUCA Many-Cores
par: Shen, Yixian, et autres
Publié: (2026)
par: Shen, Yixian, et autres
Publié: (2026)
Exploring Quantization and Mapping Synergy in Hardware-Aware Deep Neural Network Accelerators
par: Klhufek, Jan, et autres
Publié: (2024)
par: Klhufek, Jan, et autres
Publié: (2024)
FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference
par: Hooper, Coleman, et autres
Publié: (2025)
par: Hooper, Coleman, et autres
Publié: (2025)
Rescaling-Aware Training for Efficient Deployment of Deep Learning Models on Full-Integer Hardware
par: Mueller, Lion, et autres
Publié: (2025)
par: Mueller, Lion, et autres
Publié: (2025)
Hardware-Aware Fine-Tuning of Spiking Q-Networks on the SpiNNaker2 Neuromorphic Platform
par: Arfa, Sirine, et autres
Publié: (2025)
par: Arfa, Sirine, et autres
Publié: (2025)
A Hardware-Aware, Per-Layer Methodology for Post-Training Quantization of Large Language Models
par: Killian, Earl
Publié: (2026)
par: Killian, Earl
Publié: (2026)
Hardware-Aware Data and Instruction Mapping for AI Tasks: Balancing Parallelism, I/O and Memory Tradeoffs
par: Chowdhury, Md Rownak Hossain, et autres
Publié: (2025)
par: Chowdhury, Md Rownak Hossain, et autres
Publié: (2025)
LUT-DLA: Lookup Table as Efficient Extreme Low-Bit Deep Learning Accelerator
par: Li, Guoyu, et autres
Publié: (2025)
par: Li, Guoyu, et autres
Publié: (2025)
TENET: An Efficient Sparsity-Aware LUT-Centric Architecture for Ternary LLM Inference On Edge
par: Huang, Zhirui, et autres
Publié: (2025)
par: Huang, Zhirui, et autres
Publié: (2025)
Data-Rate-Aware High-Speed CNN Inference on FPGAs
par: Habermann, Tobias, et autres
Publié: (2026)
par: Habermann, Tobias, et autres
Publié: (2026)
MaRVIn: A Cross-Layer Mixed-Precision RISC-V Framework for DNN Inference, from ISA Extension to Hardware Acceleration
par: Armeniakos, Giorgos, et autres
Publié: (2025)
par: Armeniakos, Giorgos, et autres
Publié: (2025)
Context-Aware Mixture-of-Experts Inference on CXL-Enabled GPU-NDP Systems
par: Fan, Zehao, et autres
Publié: (2025)
par: Fan, Zehao, et autres
Publié: (2025)
Algorithm and Hardware Co-Design for Efficient Complex-Valued Uncertainty Estimation
par: Zhang, Zehuan, et autres
Publié: (2026)
par: Zhang, Zehuan, et autres
Publié: (2026)
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
par: Huang, Wei, et autres
Publié: (2023)
par: Huang, Wei, et autres
Publié: (2023)
Documents similaires
-
PolyLUT: Learning Piecewise Polynomials for Ultra-Low Latency FPGA LUT-based Inference
par: Andronic, Marta, et autres
Publié: (2023) -
NeuraLUT: Hiding Neural Network Density in Boolean Synthesizable Functions
par: Andronic, Marta, et autres
Publié: (2024) -
PolyLUT-Add: FPGA-based LUT Inference with Wide Inputs
par: Lou, Binglei, et autres
Publié: (2024) -
ReducedLUT: Table Decomposition with "Don't Care" Conditions
par: Cassidy, Oliver, et autres
Publié: (2024) -
ATHEENA: A Toolflow for Hardware Early-Exit Network Automation
par: Biggs, Benjamin, et autres
Publié: (2023)