H2PIPE: High throughput CNN Inference on FPGAs with High-Bandwidth Memory
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Doumet, Mario, Stan, Marius, Hall, Mathew, Betz, Vaughn |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Data-Rate-Aware High-Speed CNN Inference on FPGAs
von: Habermann, Tobias, et al.
Veröffentlicht: (2026)
von: Habermann, Tobias, et al.
Veröffentlicht: (2026)
HLSTransform: Energy-Efficient Llama 2 Inference on FPGAs Via High Level Synthesis
von: He, Andy, et al.
Veröffentlicht: (2024)
von: He, Andy, et al.
Veröffentlicht: (2024)
Enabling Long FFT Convolutions on Memory-Constrained FPGAs via Chunking
von: Wang, Peter, et al.
Veröffentlicht: (2025)
von: Wang, Peter, et al.
Veröffentlicht: (2025)
Runtime Tunable Tsetlin Machines for Edge Inference on eFPGAs
von: Rahman, Tousif, et al.
Veröffentlicht: (2025)
von: Rahman, Tousif, et al.
Veröffentlicht: (2025)
Binary Weight Multi-Bit Activation Quantization for Compute-in-Memory CNN Accelerators
von: Zhou, Wenyong, et al.
Veröffentlicht: (2025)
von: Zhou, Wenyong, et al.
Veröffentlicht: (2025)
Leveraging Highly Approximated Multipliers in DNN Inference
von: Zervakis, Georgios, et al.
Veröffentlicht: (2024)
von: Zervakis, Georgios, et al.
Veröffentlicht: (2024)
Prefill vs. Decode Bottlenecks: SRAM-Frequency Tradeoffs and the Memory-Bandwidth Ceiling
von: Atmer, Hannah, et al.
Veröffentlicht: (2025)
von: Atmer, Hannah, et al.
Veröffentlicht: (2025)
LL-GNN: Low Latency Graph Neural Networks on FPGAs for High Energy Physics
von: Que, Zhiqiang, et al.
Veröffentlicht: (2022)
von: Que, Zhiqiang, et al.
Veröffentlicht: (2022)
Dynamic Tsetlin Machine Accelerators for On-Chip Training at the Edge using FPGAs
von: Mao, Gang, et al.
Veröffentlicht: (2025)
von: Mao, Gang, et al.
Veröffentlicht: (2025)
Continuous-Flow Data-Rate-Aware CNN Inference on FPGA
von: Habermann, Tobias, et al.
Veröffentlicht: (2026)
von: Habermann, Tobias, et al.
Veröffentlicht: (2026)
Field-Programmable Gate Array Architecture for Deep Learning: Survey & Future Directions
von: Boutros, Andrew, et al.
Veröffentlicht: (2024)
von: Boutros, Andrew, et al.
Veröffentlicht: (2024)
TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs
von: Qiao, Ye, et al.
Veröffentlicht: (2025)
von: Qiao, Ye, et al.
Veröffentlicht: (2025)
TsetlinWiSARD: On-Chip Training of Weightless Neural Networks using Tsetlin Automata on FPGAs
von: Duan, Shengyu, et al.
Veröffentlicht: (2026)
von: Duan, Shengyu, et al.
Veröffentlicht: (2026)
Memory Is All You Need: An Overview of Compute-in-Memory Architectures for Accelerating Large Language Model Inference
von: Wolters, Christopher, et al.
Veröffentlicht: (2024)
von: Wolters, Christopher, et al.
Veröffentlicht: (2024)
TeLLMe v2: An Efficient End-to-End Ternary LLM Prefill and Decode Accelerator with Table-Lookup Matmul on Edge FPGAs
von: Qiao, Ye, et al.
Veröffentlicht: (2025)
von: Qiao, Ye, et al.
Veröffentlicht: (2025)
Efficient Message Passing Architecture for GCN Training on HBM-based FPGAs with Orthogonal Topology On-Chip Networks
von: Wu, Qizhe, et al.
Veröffentlicht: (2024)
von: Wu, Qizhe, et al.
Veröffentlicht: (2024)
Architectural Implications of Neural Network Inference for High Data-Rate, Low-Latency Scientific Applications
von: Weng, Olivia, et al.
Veröffentlicht: (2024)
von: Weng, Olivia, et al.
Veröffentlicht: (2024)
vMCU: Coordinated Memory Management and Kernel Optimization for DNN Inference on MCUs
von: Zheng, Size, et al.
Veröffentlicht: (2024)
von: Zheng, Size, et al.
Veröffentlicht: (2024)
CHIME: Chiplet-based Heterogeneous Near-Memory Acceleration for Edge Multimodal LLM Inference
von: Chen, Yanru, et al.
Veröffentlicht: (2025)
von: Chen, Yanru, et al.
Veröffentlicht: (2025)
Scaling Routers with In-Package Optics and High-Bandwidth Memories
von: Keslassy, Isaac, et al.
Veröffentlicht: (2026)
von: Keslassy, Isaac, et al.
Veröffentlicht: (2026)
SHIELD: A Segmented Hierarchical Memory Architecture for Energy-Efficient LLM Inference on Edge NPUs
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
A Runtime-Adaptive Transformer Neural Network Accelerator on FPGAs
von: Kabir, Ehsan, et al.
Veröffentlicht: (2024)
von: Kabir, Ehsan, et al.
Veröffentlicht: (2024)
InstantFT: An FPGA-Based Runtime Subsecond Fine-tuning of CNN Models
von: Sugiura, Keisuke, et al.
Veröffentlicht: (2025)
von: Sugiura, Keisuke, et al.
Veröffentlicht: (2025)
Revealing CNN Architectures via Side-Channel Analysis in Dataflow-based Inference Accelerators
von: Weerasena, Hansika, et al.
Veröffentlicht: (2023)
von: Weerasena, Hansika, et al.
Veröffentlicht: (2023)
FAMOUS: Flexible Accelerator for the Attention Mechanism of Transformer on UltraScale+ FPGAs
von: Kabir, Ehsan, et al.
Veröffentlicht: (2024)
von: Kabir, Ehsan, et al.
Veröffentlicht: (2024)
Inference-to-complete: A High-performance and Programmable Data-plane Co-processor for Neural-network-driven Traffic Analysis
von: Wen, Dong, et al.
Veröffentlicht: (2024)
von: Wen, Dong, et al.
Veröffentlicht: (2024)
fpgaHART: A toolflow for throughput-oriented acceleration of 3D CNNs for HAR onto FPGAs
von: Toupas, Petros, et al.
Veröffentlicht: (2023)
von: Toupas, Petros, et al.
Veröffentlicht: (2023)
Deep Inverse Design for High-Level Synthesis
von: Chang, Ping, et al.
Veröffentlicht: (2024)
von: Chang, Ping, et al.
Veröffentlicht: (2024)
Hardware/Software Co-Design of RISC-V Extensions for Accelerating Sparse DNNs on FPGAs
von: Sabih, Muhammad, et al.
Veröffentlicht: (2025)
von: Sabih, Muhammad, et al.
Veröffentlicht: (2025)
Autoformalizing Memory Specifications with Agents
von: Ernst, Jan Ole, et al.
Veröffentlicht: (2026)
von: Ernst, Jan Ole, et al.
Veröffentlicht: (2026)
DOMAC: Differentiable Optimization for High-Speed Multipliers and Multiply-Accumulators
von: Xue, Chenhao, et al.
Veröffentlicht: (2025)
von: Xue, Chenhao, et al.
Veröffentlicht: (2025)
On-Device Qwen2.5: Efficient LLM Inference with Model Compression and Hardware Acceleration
von: Xiang, Maoyang, et al.
Veröffentlicht: (2025)
von: Xiang, Maoyang, et al.
Veröffentlicht: (2025)
FLAASH: Flexible Accelerator Architecture for Sparse High-Order Tensor Contraction
von: Kulp, Gabriel, et al.
Veröffentlicht: (2024)
von: Kulp, Gabriel, et al.
Veröffentlicht: (2024)
A Hybrid Edge Classifier: Combining TinyML-Optimised CNN with RRAM-CMOS ACAM for Energy-Efficient Inference
von: Woodward, Kieran, et al.
Veröffentlicht: (2025)
von: Woodward, Kieran, et al.
Veröffentlicht: (2025)
Kernel Approximation using Analog In-Memory Computing
von: Büchel, Julian, et al.
Veröffentlicht: (2024)
von: Büchel, Julian, et al.
Veröffentlicht: (2024)
CAMformer: Associative Memory is All You Need
von: Molom-Ochir, Tergel, et al.
Veröffentlicht: (2025)
von: Molom-Ochir, Tergel, et al.
Veröffentlicht: (2025)
Leveraging Stochastic Depth Training for Adaptive Inference
von: Korol, Guilherme, et al.
Veröffentlicht: (2025)
von: Korol, Guilherme, et al.
Veröffentlicht: (2025)
HLSFactory: A Framework Empowering High-Level Synthesis Datasets for Machine Learning and Beyond
von: Abi-Karam, Stefan, et al.
Veröffentlicht: (2024)
von: Abi-Karam, Stefan, et al.
Veröffentlicht: (2024)
Efficient In-Memory Acceleration of Sparse Block Diagonal LLMs
von: de Lima, João Paulo Cardoso, et al.
Veröffentlicht: (2025)
von: de Lima, João Paulo Cardoso, et al.
Veröffentlicht: (2025)
EPIM: Efficient Processing-In-Memory Accelerators based on Epitome
von: Wang, Chenyu, et al.
Veröffentlicht: (2023)
von: Wang, Chenyu, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Data-Rate-Aware High-Speed CNN Inference on FPGAs
von: Habermann, Tobias, et al.
Veröffentlicht: (2026) -
HLSTransform: Energy-Efficient Llama 2 Inference on FPGAs Via High Level Synthesis
von: He, Andy, et al.
Veröffentlicht: (2024) -
Enabling Long FFT Convolutions on Memory-Constrained FPGAs via Chunking
von: Wang, Peter, et al.
Veröffentlicht: (2025) -
Runtime Tunable Tsetlin Machines for Edge Inference on eFPGAs
von: Rahman, Tousif, et al.
Veröffentlicht: (2025) -
Binary Weight Multi-Bit Activation Quantization for Compute-in-Memory CNN Accelerators
von: Zhou, Wenyong, et al.
Veröffentlicht: (2025)