Saved in:
| Main Authors: | Wu, Qizhe, Zhao, Letian, Gui, Yuchen, Wang, Huawen Liang Xiaotian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2411.03857 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EN-T: Optimizing Tensor Computing Engines Performance via Encoder-Based Methodology
by: Wu, Qizhe, et al.
Published: (2024)
by: Wu, Qizhe, et al.
Published: (2024)
Exploring the Performance Improvement of Tensor Processing Engines through Transformation in the Bit-weight Dimension of MACs
by: Wu, Qizhe, et al.
Published: (2025)
by: Wu, Qizhe, et al.
Published: (2025)
Dynamic Tsetlin Machine Accelerators for On-Chip Training at the Edge using FPGAs
by: Mao, Gang, et al.
Published: (2025)
by: Mao, Gang, et al.
Published: (2025)
TsetlinWiSARD: On-Chip Training of Weightless Neural Networks using Tsetlin Automata on FPGAs
by: Duan, Shengyu, et al.
Published: (2026)
by: Duan, Shengyu, et al.
Published: (2026)
TAPA-CS: Enabling Scalable Accelerator Design on Distributed HBM-FPGAs
by: Prakriya, Neha, et al.
Published: (2023)
by: Prakriya, Neha, et al.
Published: (2023)
LlamaF: An Efficient Llama2 Architecture Accelerator on Embedded FPGAs
by: Xu, Han, et al.
Published: (2024)
by: Xu, Han, et al.
Published: (2024)
GCN-ABFT: Low-Cost Online Error Checking for Graph Convolutional Networks
by: Peltekis, Christodoulos, et al.
Published: (2024)
by: Peltekis, Christodoulos, et al.
Published: (2024)
Energy Efficient LSTM Accelerators for Embedded FPGAs through Parameterised Architecture Design
by: Qian, Chao, et al.
Published: (2026)
by: Qian, Chao, et al.
Published: (2026)
TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs
by: Qiao, Ye, et al.
Published: (2025)
by: Qiao, Ye, et al.
Published: (2025)
Read Disturbance in High Bandwidth Memory: A Detailed Experimental Study on HBM2 DRAM Chips
by: Olgun, Ataberk, et al.
Published: (2023)
by: Olgun, Ataberk, et al.
Published: (2023)
SMOF: Streaming Modern CNNs on FPGAs with Smart Off-Chip Eviction
by: Toupas, Petros, et al.
Published: (2024)
by: Toupas, Petros, et al.
Published: (2024)
AutoFlows++: Hierarchical Message Flow Mining for System on Chip Designs
by: Nadimi, Bardia, et al.
Published: (2026)
by: Nadimi, Bardia, et al.
Published: (2026)
Data-Rate-Aware High-Speed CNN Inference on FPGAs
by: Habermann, Tobias, et al.
Published: (2026)
by: Habermann, Tobias, et al.
Published: (2026)
A Survey on LUT-based Deep Neural Networks Implemented in FPGAs
by: Guo, Zeyu
Published: (2025)
by: Guo, Zeyu
Published: (2025)
Leveraging FPGAs for Homomorphic Matrix-Vector Multiplication in Oblivious Message Retrieval
by: Bosworth, Grant, et al.
Published: (2025)
by: Bosworth, Grant, et al.
Published: (2025)
A Runtime-Adaptive Transformer Neural Network Accelerator on FPGAs
by: Kabir, Ehsan, et al.
Published: (2024)
by: Kabir, Ehsan, et al.
Published: (2024)
Enabling Long FFT Convolutions on Memory-Constrained FPGAs via Chunking
by: Wang, Peter, et al.
Published: (2025)
by: Wang, Peter, et al.
Published: (2025)
Making Strong Error-Correcting Codes Work Effectively for HBM in AI Inference
by: Xie, Rui, et al.
Published: (2025)
by: Xie, Rui, et al.
Published: (2025)
TeLLMe v2: An Efficient End-to-End Ternary LLM Prefill and Decode Accelerator with Table-Lookup Matmul on Edge FPGAs
by: Qiao, Ye, et al.
Published: (2025)
by: Qiao, Ye, et al.
Published: (2025)
Breaking the HBM Bit Cost Barrier: Domain-Specific ECC for AI Inference Infrastructure
by: Xie, Rui, et al.
Published: (2025)
by: Xie, Rui, et al.
Published: (2025)
Experimental Demonstration of an Optical Neural PDE Solver via On-Chip PINN Training
by: Zhao, Yequan, et al.
Published: (2025)
by: Zhao, Yequan, et al.
Published: (2025)
H2PIPE: High throughput CNN Inference on FPGAs with High-Bandwidth Memory
by: Doumet, Mario, et al.
Published: (2024)
by: Doumet, Mario, et al.
Published: (2024)
JEDI-linear: Fast and Efficient Graph Neural Networks for Jet Tagging on FPGAs
by: Que, Zhiqiang, et al.
Published: (2025)
by: Que, Zhiqiang, et al.
Published: (2025)
Accelerating Boolean Constraint Propagation for Efficient SAT-Solving on FPGAs
by: Govindasamy, Hariprasadh, et al.
Published: (2024)
by: Govindasamy, Hariprasadh, et al.
Published: (2024)
Efficient Approaches for GEMM Acceleration on Leading AI-Optimized FPGAs
by: Taka, Endri, et al.
Published: (2024)
by: Taka, Endri, et al.
Published: (2024)
HLSTransform: Energy-Efficient Llama 2 Inference on FPGAs Via High Level Synthesis
by: He, Andy, et al.
Published: (2024)
by: He, Andy, et al.
Published: (2024)
LL-GNN: Low Latency Graph Neural Networks on FPGAs for High Energy Physics
by: Que, Zhiqiang, et al.
Published: (2022)
by: Que, Zhiqiang, et al.
Published: (2022)
QiMeng: Fully Automated Hardware and Software Design for Processor Chip
by: Zhang, Rui, et al.
Published: (2025)
by: Zhang, Rui, et al.
Published: (2025)
FedChip: Federated LLM for Artificial Intelligence Accelerator Chip Design
by: Nazzal, Mahmoud, et al.
Published: (2025)
by: Nazzal, Mahmoud, et al.
Published: (2025)
SparsePixels: Efficient Convolution for Sparse Data on FPGAs
by: Tsoi, Ho Fung, et al.
Published: (2025)
by: Tsoi, Ho Fung, et al.
Published: (2025)
Small Logic-based Multipliers with Incomplete Sub-Multipliers for FPGAs
by: Böttcher, Andreas, et al.
Published: (2024)
by: Böttcher, Andreas, et al.
Published: (2024)
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
by: Huang, Wei, et al.
Published: (2023)
by: Huang, Wei, et al.
Published: (2023)
ONNX-to-Hardware Design Flow for Adaptive Neural-Network Inference on FPGAs
by: Manca, Federico, et al.
Published: (2024)
by: Manca, Federico, et al.
Published: (2024)
Low-latency D-MIMO Localization using Distributed Scalable Message-Passing Algorithm
by: Iancu, Dumitra, et al.
Published: (2025)
by: Iancu, Dumitra, et al.
Published: (2025)
Hardware-Efficient Accurate 4-bit Multiplier for Xilinx 7 Series FPGAs
by: Kida, Misaki, et al.
Published: (2025)
by: Kida, Misaki, et al.
Published: (2025)
An Energy-Efficient Artefact Detection Accelerator on FPGAs for Hyper-Spectral Satellite Imagery
by: Castelino, Cornell, et al.
Published: (2024)
by: Castelino, Cornell, et al.
Published: (2024)
SWAT: Scalable and Efficient Window Attention-based Transformers Acceleration on FPGAs
by: Bai, Zhenyu, et al.
Published: (2024)
by: Bai, Zhenyu, et al.
Published: (2024)
Memory-efficient Sketch Acceleration for Handling Large Network Flows on FPGAs
by: Han, Zhaoyang, et al.
Published: (2025)
by: Han, Zhaoyang, et al.
Published: (2025)
Explainable Fuzzy Neural Network with Multi-Fidelity Reinforcement Learning for Micro-Architecture Design Space Exploration
by: Fan, Hanwei, et al.
Published: (2024)
by: Fan, Hanwei, et al.
Published: (2024)
TokenStack: A Heterogeneous HBM-PIM Architecture and Runtime for Efficient LLM Inference
by: Li, Zhuoran, et al.
Published: (2026)
by: Li, Zhuoran, et al.
Published: (2026)
Similar Items
-
EN-T: Optimizing Tensor Computing Engines Performance via Encoder-Based Methodology
by: Wu, Qizhe, et al.
Published: (2024) -
Exploring the Performance Improvement of Tensor Processing Engines through Transformation in the Bit-weight Dimension of MACs
by: Wu, Qizhe, et al.
Published: (2025) -
Dynamic Tsetlin Machine Accelerators for On-Chip Training at the Edge using FPGAs
by: Mao, Gang, et al.
Published: (2025) -
TsetlinWiSARD: On-Chip Training of Weightless Neural Networks using Tsetlin Automata on FPGAs
by: Duan, Shengyu, et al.
Published: (2026) -
TAPA-CS: Enabling Scalable Accelerator Design on Distributed HBM-FPGAs
by: Prakriya, Neha, et al.
Published: (2023)