Design and Implementation of an FPGA-Based Hardware Accelerator for Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Richie, Chen, Sicheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dataflow & Tiling Strategies in Edge-AI FPGA Accelerators: A Comprehensive Literature Review
by: Li, Richie
Published: (2025)
by: Li, Richie
Published: (2025)
RISCBench: Benchmarking RISC-V Orchestration Efficiency in FPGA and FPGA-Like Computing Engines
by: Ojika, Dave, et al.
Published: (2025)
by: Ojika, Dave, et al.
Published: (2025)
RV-IM100: Quantifying ISA Extension, Datapath Width, and Pipeline Depth Trade-offs in RISC-V Microarchitectures
by: Kang, Hyunwoo
Published: (2026)
by: Kang, Hyunwoo
Published: (2026)
RISC-V Based TinyML Accelerator for Depthwise Separable Convolutions in Edge AI
by: Yildirim, Muhammed, et al.
Published: (2025)
by: Yildirim, Muhammed, et al.
Published: (2025)
Lean Attention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers
by: Sanovar, Rya, et al.
Published: (2024)
by: Sanovar, Rya, et al.
Published: (2024)
TOM: A Ternary Read-only Memory Accelerator for LLM-powered Edge Intelligence
by: Guan, Hongyi, et al.
Published: (2026)
by: Guan, Hongyi, et al.
Published: (2026)
Biological Intuition on Digital Hardware: An RTL Implementation of Poisson-Encoded SNNs for Static Image Classification
by: Das, Debabrata, et al.
Published: (2026)
by: Das, Debabrata, et al.
Published: (2026)
NeuroPDE: A Neuromorphic PDE Solver Based on Spintronic and Ferroelectric Devices
by: Fu, Siqing, et al.
Published: (2025)
by: Fu, Siqing, et al.
Published: (2025)
basic_RV32s: An Open-Source Microarchitectural Roadmap for RISC-V RV32I
by: Kang, Hyun Woo, et al.
Published: (2025)
by: Kang, Hyun Woo, et al.
Published: (2025)
DSPE: An Energy-Efficient Edge Processor for DeepSeek Inference with MerkleTree-based Incremental Pruning, Multi-Stage Boothing Lookup and Dynamic Adaptive Posit Processing
by: Zhang, Yuhan, et al.
Published: (2026)
by: Zhang, Yuhan, et al.
Published: (2026)
Not All Thoughts Need HBM: Semantics-Aware Memory Hierarchy for LLM Reasoning
by: Yuan, Aojie, et al.
Published: (2026)
by: Yuan, Aojie, et al.
Published: (2026)
ECOLogic: Enabling Circular, Obfuscated, and Adaptive Logic via eFPGA-Augmented SoCs
by: Tashdid, Ishraq, et al.
Published: (2025)
by: Tashdid, Ishraq, et al.
Published: (2025)
D-com: Accelerating Iterative Processing to Enable Low-rank Decomposition of Activations
by: Tahmasebi, Faraz, et al.
Published: (2025)
by: Tahmasebi, Faraz, et al.
Published: (2025)
TokenStack: A Heterogeneous HBM-PIM Architecture and Runtime for Efficient LLM Inference
by: Li, Zhuoran, et al.
Published: (2026)
by: Li, Zhuoran, et al.
Published: (2026)
A flexible framework for early power and timing comparison of time-multiplexed CGRA kernel executions
by: Aspros, Maxime Henri, et al.
Published: (2025)
by: Aspros, Maxime Henri, et al.
Published: (2025)
A Survey on Hardware Accelerators for Large Language Models
by: Kachris, Christoforos
Published: (2024)
by: Kachris, Christoforos
Published: (2024)
Time Domain Near Memory Computing Engine
by: Antal, Sarthak, et al.
Published: (2026)
by: Antal, Sarthak, et al.
Published: (2026)
Nonvolatile Charge-Domain Attention with HZO Ferroelectric Capacitors: A Simulation-Based Device-to-System Evaluation
by: Abouagour, Faris
Published: (2026)
by: Abouagour, Faris
Published: (2026)
Factor Machine: Mixed-signal Architecture for Fine-Grained Graph-Based Computing
by: Dudek, Piotr
Published: (2024)
by: Dudek, Piotr
Published: (2024)
CGRA4ML: A Hardware/Software Framework to Implement Neural Networks for Scientific Edge Computing
by: Abarajithan, G, et al.
Published: (2024)
by: Abarajithan, G, et al.
Published: (2024)
InterPUF: Distributed Authentication via Physically Unclonable Functions and Multi-party Computation for Reconfigurable Interposers
by: Tashdid, Ishraq, et al.
Published: (2026)
by: Tashdid, Ishraq, et al.
Published: (2026)
LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM Inference
by: Mo, Zhiwen, et al.
Published: (2024)
by: Mo, Zhiwen, et al.
Published: (2024)
Minimal Neuron Circuits -- Part I: Resonators
by: Nabil, Amr, et al.
Published: (2025)
by: Nabil, Amr, et al.
Published: (2025)
Minimal Neuron Circuits: Bursters
by: Nabil, Amr, et al.
Published: (2025)
by: Nabil, Amr, et al.
Published: (2025)
Ten-Four: An Open-Source Fused Dot Product Unit for Mixed-Precision GPGPU Tensor Cores
by: Rout, Nikhil, et al.
Published: (2025)
by: Rout, Nikhil, et al.
Published: (2025)
A 65 nm Bayesian Neural Network Accelerator with 360 fJ/Sample In-Word GRNG for AI Uncertainty Estimation
by: Enciso, Zephan M., et al.
Published: (2025)
by: Enciso, Zephan M., et al.
Published: (2025)
GainSight: A Unified Framework for Data Lifetime Profiling and Heterogeneous Memory Composition
by: Li, Peijing, et al.
Published: (2025)
by: Li, Peijing, et al.
Published: (2025)
FOCUS: DLLMs Know How to Tame Their Compute Bound
by: Liang, Kaihua, et al.
Published: (2026)
by: Liang, Kaihua, et al.
Published: (2026)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
by: Jo, Myeong Jun
Published: (2026)
by: Jo, Myeong Jun
Published: (2026)
RayFlex: An Open-Source RTL Implementation of the Hardware Ray Tracer Datapath
by: Shen, Fangjia, et al.
Published: (2024)
by: Shen, Fangjia, et al.
Published: (2024)
Lincoln AI Computing Survey (LAICS) and Trends
by: Reuther, Albert, et al.
Published: (2025)
by: Reuther, Albert, et al.
Published: (2025)
FlexiBit: Fully Flexible Precision Bit-parallel Accelerator Architecture for Arbitrary Mixed Precision AI
by: Tahmasebi, Faraz, et al.
Published: (2024)
by: Tahmasebi, Faraz, et al.
Published: (2024)
An Integrated UVM-TLM Co-Simulation Framework for RISC-V Functional Verification and Performance Evaluation
by: Qiu, Ruizhi, et al.
Published: (2025)
by: Qiu, Ruizhi, et al.
Published: (2025)
Toward a Universal GPU Instruction Set Architecture: A Cross-Vendor Analysis of Hardware-Invariant Computational Primitives in Parallel Processors
by: Abraham, Ojima, et al.
Published: (2026)
by: Abraham, Ojima, et al.
Published: (2026)
A Per-Access Upper Bound for Shared-Resource Interference in Direct-Mapped Multicore Architectures
by: Pedroni, Felipe T.
Published: (2026)
by: Pedroni, Felipe T.
Published: (2026)
Taming Wild Branches: Overcoming Hard-to-Predict Branches using the Bullseye Predictor
by: Behrendt, Emet, et al.
Published: (2025)
by: Behrendt, Emet, et al.
Published: (2025)
Ember: A Compiler for Efficient Embedding Operations on Decoupled Access-Execute Architectures
by: Siracusa, Marco, et al.
Published: (2025)
by: Siracusa, Marco, et al.
Published: (2025)
An All-digital 8.6-nJ/Frame 65-nm Tsetlin Machine Image Classification Accelerator
by: Tunheim, Svein Anders, et al.
Published: (2025)
by: Tunheim, Svein Anders, et al.
Published: (2025)
Accelerating Precise End-to-End Simulation: Latency-Sensitive Many-core System Modeling
by: Li, Yinrong, et al.
Published: (2026)
by: Li, Yinrong, et al.
Published: (2026)
Ultra Memory-Efficient On-FPGA Training of Transformers via Tensor-Compressed Optimization
by: Tian, Jiayi, et al.
Published: (2025)
by: Tian, Jiayi, et al.
Published: (2025)
Similar Items
-
Dataflow & Tiling Strategies in Edge-AI FPGA Accelerators: A Comprehensive Literature Review
by: Li, Richie
Published: (2025) -
RISCBench: Benchmarking RISC-V Orchestration Efficiency in FPGA and FPGA-Like Computing Engines
by: Ojika, Dave, et al.
Published: (2025) -
RV-IM100: Quantifying ISA Extension, Datapath Width, and Pipeline Depth Trade-offs in RISC-V Microarchitectures
by: Kang, Hyunwoo
Published: (2026) -
RISC-V Based TinyML Accelerator for Depthwise Separable Convolutions in Edge AI
by: Yildirim, Muhammed, et al.
Published: (2025) -
Lean Attention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers
by: Sanovar, Rya, et al.
Published: (2024)