Continuous-Flow Data-Rate-Aware CNN Inference on FPGA
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Habermann, Tobias, Mecik, Michael, Wang, Zhenyu, Vera, César David, Kumm, Martin, Garrido, Mario |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Data-Rate-Aware High-Speed CNN Inference on FPGAs
von: Habermann, Tobias, et al.
Veröffentlicht: (2026)
von: Habermann, Tobias, et al.
Veröffentlicht: (2026)
Implementation and Analysis of Thermometer Encoding in DWN FPGA Accelerators
von: Mecik, Michael, et al.
Veröffentlicht: (2025)
von: Mecik, Michael, et al.
Veröffentlicht: (2025)
TRINE: A Token-Aware, Runtime-Adaptive FPGA Inference Engine for Multimodal AI
von: Oh, Hyunwoo, et al.
Veröffentlicht: (2026)
von: Oh, Hyunwoo, et al.
Veröffentlicht: (2026)
PolyLUT-Add: FPGA-based LUT Inference with Wide Inputs
von: Lou, Binglei, et al.
Veröffentlicht: (2024)
von: Lou, Binglei, et al.
Veröffentlicht: (2024)
LUTMUL: Exceed Conventional FPGA Roofline Limit by LUT-based Efficient Multiplication for Neural Network Inference
von: Xie, Yanyue, et al.
Veröffentlicht: (2024)
von: Xie, Yanyue, et al.
Veröffentlicht: (2024)
At the Edge of the Heart: ULP FPGA-Based CNN for On-Device Cardiac Feature Extraction in Smart Health Sensors for Astronauts
von: Rahman, Kazi Mohammad Abidur, et al.
Veröffentlicht: (2026)
von: Rahman, Kazi Mohammad Abidur, et al.
Veröffentlicht: (2026)
fSEAD: a Composable FPGA-based Streaming Ensemble Anomaly Detection Library
von: Lou, Binglei, et al.
Veröffentlicht: (2024)
von: Lou, Binglei, et al.
Veröffentlicht: (2024)
A Hybrid Edge Classifier: Combining TinyML-Optimised CNN with RRAM-CMOS ACAM for Energy-Efficient Inference
von: Woodward, Kieran, et al.
Veröffentlicht: (2025)
von: Woodward, Kieran, et al.
Veröffentlicht: (2025)
Embedded FPGA Acceleration of Brain-Like Neural Networks: Online Learning to Scalable Inference
von: Hafiz, Muhammad Ihsan Al, et al.
Veröffentlicht: (2025)
von: Hafiz, Muhammad Ihsan Al, et al.
Veröffentlicht: (2025)
IMAGINE: An 8-to-1b 22nm FD-SOI Compute-In-Memory CNN Accelerator With an End-to-End Analog Charge-Based 0.15-8POPS/W Macro Featuring Distribution-Aware Data Reshaping
von: Kneip, Adrian, et al.
Veröffentlicht: (2024)
von: Kneip, Adrian, et al.
Veröffentlicht: (2024)
FPGA Divide-and-Conquer Placement using Deep Reinforcement Learning
von: Wang, Shang, et al.
Veröffentlicht: (2024)
von: Wang, Shang, et al.
Veröffentlicht: (2024)
Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inference
von: Chen, Hongzheng, et al.
Veröffentlicht: (2023)
von: Chen, Hongzheng, et al.
Veröffentlicht: (2023)
Fast, Scalable, Energy-Efficient Non-element-wise Matrix Multiplication on FPGA
von: Zhu, Xuqi, et al.
Veröffentlicht: (2024)
von: Zhu, Xuqi, et al.
Veröffentlicht: (2024)
ALADIN: Accuracy-Latency-Aware Design-space Inference Analysis for Embedded AI Accelerators
von: Baldi, T., et al.
Veröffentlicht: (2026)
von: Baldi, T., et al.
Veröffentlicht: (2026)
Idle is the New Sleep: Configuration-Aware Alternative to Powering Off FPGA-Based DL Accelerators During Inactivity
von: Qian, Chao, et al.
Veröffentlicht: (2024)
von: Qian, Chao, et al.
Veröffentlicht: (2024)
HAPM -- Hardware Aware Pruning Method for CNN hardware accelerators in resource constrained devices
von: Peccia, Federico Nicolas, et al.
Veröffentlicht: (2024)
von: Peccia, Federico Nicolas, et al.
Veröffentlicht: (2024)
rule4ml: An Open-Source Tool for Resource Utilization and Latency Estimation for ML Models on FPGA
von: Rahimifar, Mohammad Mehdi, et al.
Veröffentlicht: (2024)
von: Rahimifar, Mohammad Mehdi, et al.
Veröffentlicht: (2024)
Small Logic-based Multipliers with Incomplete Sub-Multipliers for FPGAs
von: Böttcher, Andreas, et al.
Veröffentlicht: (2024)
von: Böttcher, Andreas, et al.
Veröffentlicht: (2024)
Multiplier Design Addressing Area-Delay Trade-offs by using DSP and Logic resources on FPGAs
von: Böttcher, Andreas, et al.
Veröffentlicht: (2024)
von: Böttcher, Andreas, et al.
Veröffentlicht: (2024)
Challenges and Research Directions for Large Language Model Inference Hardware
von: Ma, Xiaoyu, et al.
Veröffentlicht: (2026)
von: Ma, Xiaoyu, et al.
Veröffentlicht: (2026)
ProTEA: Programmable Transformer Encoder Acceleration on FPGA
von: Kabir, Ehsan, et al.
Veröffentlicht: (2024)
von: Kabir, Ehsan, et al.
Veröffentlicht: (2024)
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format
von: Fang, Chao, et al.
Veröffentlicht: (2024)
von: Fang, Chao, et al.
Veröffentlicht: (2024)
FlowPlace: Flow Matching for Chip Placement
von: Xie, Peng, et al.
Veröffentlicht: (2026)
von: Xie, Peng, et al.
Veröffentlicht: (2026)
NSFlow: An End-to-End FPGA Framework with Scalable Dataflow Architecture for Neuro-Symbolic AI
von: Yang, Hanchen, et al.
Veröffentlicht: (2025)
von: Yang, Hanchen, et al.
Veröffentlicht: (2025)
FPGA Resource-aware Structured Pruning for Real-Time Neural Networks
von: Ramhorst, Benjamin, et al.
Veröffentlicht: (2023)
von: Ramhorst, Benjamin, et al.
Veröffentlicht: (2023)
FPGA-Based Neural Network Accelerators for Space Applications: A Survey
von: Antunes, Pedro, et al.
Veröffentlicht: (2025)
von: Antunes, Pedro, et al.
Veröffentlicht: (2025)
HiFloat4 Format for Language Model Inference
von: Luo, Yuanyong, et al.
Veröffentlicht: (2026)
von: Luo, Yuanyong, et al.
Veröffentlicht: (2026)
FastMamba: A High-Speed and Efficient Mamba Accelerator on FPGA with Accurate Quantization
von: Wang, Aotao, et al.
Veröffentlicht: (2025)
von: Wang, Aotao, et al.
Veröffentlicht: (2025)
Architectural Design and Performance Analysis of FPGA based AI Accelerators: A Comprehensive Review
von: Chatterjee, Soumita, et al.
Veröffentlicht: (2026)
von: Chatterjee, Soumita, et al.
Veröffentlicht: (2026)
Hardware-Efficient FPGA Implementation of Sigmoid Function Using Mixed-Radix Hyperbolic Rotation CORDIC
von: Panchal, Chintan, et al.
Veröffentlicht: (2026)
von: Panchal, Chintan, et al.
Veröffentlicht: (2026)
FlightLLM: Efficient Large Language Model Inference with a Complete Mapping Flow on FPGAs
von: Zeng, Shulin, et al.
Veröffentlicht: (2024)
von: Zeng, Shulin, et al.
Veröffentlicht: (2024)
Efficient Deployment of CNN Models on Multiple In-Memory Computing Units
von: Bougioukou, Eleni, et al.
Veröffentlicht: (2025)
von: Bougioukou, Eleni, et al.
Veröffentlicht: (2025)
Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference
von: Skliar, Andrii, et al.
Veröffentlicht: (2024)
von: Skliar, Andrii, et al.
Veröffentlicht: (2024)
Runtime Tunable Tsetlin Machines for Edge Inference on eFPGAs
von: Rahman, Tousif, et al.
Veröffentlicht: (2025)
von: Rahman, Tousif, et al.
Veröffentlicht: (2025)
Heterogeneous SoC Integrating an Open-Source Recurrent SNN Accelerator for Neuromorphic Edge Computing on FPGA
von: Barocci, Michelangelo, et al.
Veröffentlicht: (2026)
von: Barocci, Michelangelo, et al.
Veröffentlicht: (2026)
Optimizing Neural Networks with Learnable Non-Linear Activation Functions via Lookup-Based FPGA Acceleration
von: Yin, Mengyuan, et al.
Veröffentlicht: (2025)
von: Yin, Mengyuan, et al.
Veröffentlicht: (2025)
Towards Cheaper Inference in Deep Networks with Lower Bit-Width Accumulators
von: Blumenfeld, Yaniv, et al.
Veröffentlicht: (2024)
von: Blumenfeld, Yaniv, et al.
Veröffentlicht: (2024)
MixPE: Quantization and Hardware Co-design for Efficient LLM Inference
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
FINN-GL: Generalized Mixed-Precision Extensions for FPGA-Accelerated LSTMs
von: Khandelwal, Shashwat, et al.
Veröffentlicht: (2025)
von: Khandelwal, Shashwat, et al.
Veröffentlicht: (2025)
Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns
von: Bambhaniya, Abhimanyu, et al.
Veröffentlicht: (2026)
von: Bambhaniya, Abhimanyu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Data-Rate-Aware High-Speed CNN Inference on FPGAs
von: Habermann, Tobias, et al.
Veröffentlicht: (2026) -
Implementation and Analysis of Thermometer Encoding in DWN FPGA Accelerators
von: Mecik, Michael, et al.
Veröffentlicht: (2025) -
TRINE: A Token-Aware, Runtime-Adaptive FPGA Inference Engine for Multimodal AI
von: Oh, Hyunwoo, et al.
Veröffentlicht: (2026) -
PolyLUT-Add: FPGA-based LUT Inference with Wide Inputs
von: Lou, Binglei, et al.
Veröffentlicht: (2024) -
LUTMUL: Exceed Conventional FPGA Roofline Limit by LUT-based Efficient Multiplication for Neural Network Inference
von: Xie, Yanyue, et al.
Veröffentlicht: (2024)