Data-Rate-Aware High-Speed CNN Inference on FPGAs
Fuente:
arXiv
Saved in:
| Main Authors: | Habermann, Tobias, Kumm, Martin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Continuous-Flow Data-Rate-Aware CNN Inference on FPGA
by: Habermann, Tobias, et al.
Published: (2026)
by: Habermann, Tobias, et al.
Published: (2026)
H2PIPE: High throughput CNN Inference on FPGAs with High-Bandwidth Memory
by: Doumet, Mario, et al.
Published: (2024)
by: Doumet, Mario, et al.
Published: (2024)
Small Logic-based Multipliers with Incomplete Sub-Multipliers for FPGAs
by: Böttcher, Andreas, et al.
Published: (2024)
by: Böttcher, Andreas, et al.
Published: (2024)
Multiplier Design Addressing Area-Delay Trade-offs by using DSP and Logic resources on FPGAs
by: Böttcher, Andreas, et al.
Published: (2024)
by: Böttcher, Andreas, et al.
Published: (2024)
Architectural Implications of Neural Network Inference for High Data-Rate, Low-Latency Scientific Applications
by: Weng, Olivia, et al.
Published: (2024)
by: Weng, Olivia, et al.
Published: (2024)
HLSTransform: Energy-Efficient Llama 2 Inference on FPGAs Via High Level Synthesis
by: He, Andy, et al.
Published: (2024)
by: He, Andy, et al.
Published: (2024)
Runtime Tunable Tsetlin Machines for Edge Inference on eFPGAs
by: Rahman, Tousif, et al.
Published: (2025)
by: Rahman, Tousif, et al.
Published: (2025)
Implementation and Analysis of Thermometer Encoding in DWN FPGA Accelerators
by: Mecik, Michael, et al.
Published: (2025)
by: Mecik, Michael, et al.
Published: (2025)
Enabling Long FFT Convolutions on Memory-Constrained FPGAs via Chunking
by: Wang, Peter, et al.
Published: (2025)
by: Wang, Peter, et al.
Published: (2025)
Dynamic Tsetlin Machine Accelerators for On-Chip Training at the Edge using FPGAs
by: Mao, Gang, et al.
Published: (2025)
by: Mao, Gang, et al.
Published: (2025)
TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs
by: Qiao, Ye, et al.
Published: (2025)
by: Qiao, Ye, et al.
Published: (2025)
DOMAC: Differentiable Optimization for High-Speed Multipliers and Multiply-Accumulators
by: Xue, Chenhao, et al.
Published: (2025)
by: Xue, Chenhao, et al.
Published: (2025)
A Data-Driven Approach to Dataflow-Aware Online Scheduling for Graph Neural Network Inference
by: Puigdemont, Pol, et al.
Published: (2024)
by: Puigdemont, Pol, et al.
Published: (2024)
TsetlinWiSARD: On-Chip Training of Weightless Neural Networks using Tsetlin Automata on FPGAs
by: Duan, Shengyu, et al.
Published: (2026)
by: Duan, Shengyu, et al.
Published: (2026)
Efficient Message Passing Architecture for GCN Training on HBM-based FPGAs with Orthogonal Topology On-Chip Networks
by: Wu, Qizhe, et al.
Published: (2024)
by: Wu, Qizhe, et al.
Published: (2024)
Efficient and Mathematically Robust Operations for Certified Neural Networks Inference
by: Geyer, Fabien, et al.
Published: (2024)
by: Geyer, Fabien, et al.
Published: (2024)
LL-GNN: Low Latency Graph Neural Networks on FPGAs for High Energy Physics
by: Que, Zhiqiang, et al.
Published: (2022)
by: Que, Zhiqiang, et al.
Published: (2022)
TeLLMe v2: An Efficient End-to-End Ternary LLM Prefill and Decode Accelerator with Table-Lookup Matmul on Edge FPGAs
by: Qiao, Ye, et al.
Published: (2025)
by: Qiao, Ye, et al.
Published: (2025)
Leveraging Highly Approximated Multipliers in DNN Inference
by: Zervakis, Georgios, et al.
Published: (2024)
by: Zervakis, Georgios, et al.
Published: (2024)
SigmaQuant: Hardware-Aware Heterogeneous Quantization Method for Edge DNN Inference
by: Liu, Qunyou, et al.
Published: (2026)
by: Liu, Qunyou, et al.
Published: (2026)
Context-Aware Mixture-of-Experts Inference on CXL-Enabled GPU-NDP Systems
by: Fan, Zehao, et al.
Published: (2025)
by: Fan, Zehao, et al.
Published: (2025)
PolyLUT: Ultra-low Latency Polynomial Inference with Hardware-Aware Structured Pruning
by: Andronic, Marta, et al.
Published: (2025)
by: Andronic, Marta, et al.
Published: (2025)
Active Imitation Learning for Thermal- and Kernel-Aware LFM Inference on 3D S-NUCA Many-Cores
by: Shen, Yixian, et al.
Published: (2026)
by: Shen, Yixian, et al.
Published: (2026)
A Runtime-Adaptive Transformer Neural Network Accelerator on FPGAs
by: Kabir, Ehsan, et al.
Published: (2024)
by: Kabir, Ehsan, et al.
Published: (2024)
InstantFT: An FPGA-Based Runtime Subsecond Fine-tuning of CNN Models
by: Sugiura, Keisuke, et al.
Published: (2025)
by: Sugiura, Keisuke, et al.
Published: (2025)
Binary Weight Multi-Bit Activation Quantization for Compute-in-Memory CNN Accelerators
by: Zhou, Wenyong, et al.
Published: (2025)
by: Zhou, Wenyong, et al.
Published: (2025)
SparsePixels: Efficient Convolution for Sparse Data on FPGAs
by: Tsoi, Ho Fung, et al.
Published: (2025)
by: Tsoi, Ho Fung, et al.
Published: (2025)
Revealing CNN Architectures via Side-Channel Analysis in Dataflow-based Inference Accelerators
by: Weerasena, Hansika, et al.
Published: (2023)
by: Weerasena, Hansika, et al.
Published: (2023)
FAMOUS: Flexible Accelerator for the Attention Mechanism of Transformer on UltraScale+ FPGAs
by: Kabir, Ehsan, et al.
Published: (2024)
by: Kabir, Ehsan, et al.
Published: (2024)
Inference-to-complete: A High-performance and Programmable Data-plane Co-processor for Neural-network-driven Traffic Analysis
by: Wen, Dong, et al.
Published: (2024)
by: Wen, Dong, et al.
Published: (2024)
Hardware/Software Co-Design of RISC-V Extensions for Accelerating Sparse DNNs on FPGAs
by: Sabih, Muhammad, et al.
Published: (2025)
by: Sabih, Muhammad, et al.
Published: (2025)
Hardware-Aware Data and Instruction Mapping for AI Tasks: Balancing Parallelism, I/O and Memory Tradeoffs
by: Chowdhury, Md Rownak Hossain, et al.
Published: (2025)
by: Chowdhury, Md Rownak Hossain, et al.
Published: (2025)
A Hybrid Edge Classifier: Combining TinyML-Optimised CNN with RRAM-CMOS ACAM for Energy-Efficient Inference
by: Woodward, Kieran, et al.
Published: (2025)
by: Woodward, Kieran, et al.
Published: (2025)
Leveraging Stochastic Depth Training for Adaptive Inference
by: Korol, Guilherme, et al.
Published: (2025)
by: Korol, Guilherme, et al.
Published: (2025)
KLLM: Fast LLM Inference with K-Means Quantization
by: Wu, Xueying, et al.
Published: (2025)
by: Wu, Xueying, et al.
Published: (2025)
DPUConfig: Optimizing ML Inference in FPGAs Using Reinforcement Learning
by: Patras, Alexandros, et al.
Published: (2026)
by: Patras, Alexandros, et al.
Published: (2026)
Interconnect-Aware Logic Resynthesis for Multi-Die FPGAs
by: Wang, Xiaoke, et al.
Published: (2026)
by: Wang, Xiaoke, et al.
Published: (2026)
FORTALESA: Fault-Tolerant Reconfigurable Systolic Array for DNN Inference
by: Cherezova, Natalia, et al.
Published: (2025)
by: Cherezova, Natalia, et al.
Published: (2025)
InTAR: Inter-Task Auto-Reconfigurable Accelerator Design for High Data Volume Variation in DNNs
by: He, Zifan, et al.
Published: (2025)
by: He, Zifan, et al.
Published: (2025)
VeriCache: Turning Lossy KV Cache into Lossless LLM Inference
by: Yao, Jiayi, et al.
Published: (2026)
by: Yao, Jiayi, et al.
Published: (2026)
Similar Items
-
Continuous-Flow Data-Rate-Aware CNN Inference on FPGA
by: Habermann, Tobias, et al.
Published: (2026) -
H2PIPE: High throughput CNN Inference on FPGAs with High-Bandwidth Memory
by: Doumet, Mario, et al.
Published: (2024) -
Small Logic-based Multipliers with Incomplete Sub-Multipliers for FPGAs
by: Böttcher, Andreas, et al.
Published: (2024) -
Multiplier Design Addressing Area-Delay Trade-offs by using DSP and Logic resources on FPGAs
by: Böttcher, Andreas, et al.
Published: (2024) -
Architectural Implications of Neural Network Inference for High Data-Rate, Low-Latency Scientific Applications
by: Weng, Olivia, et al.
Published: (2024)