Dataflow & Tiling Strategies in Edge-AI FPGA Accelerators: A Comprehensive Literature Review
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Li, Richie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Design and Implementation of an FPGA-Based Hardware Accelerator for Transformer
von: Li, Richie, et al.
Veröffentlicht: (2025)
von: Li, Richie, et al.
Veröffentlicht: (2025)
TOM: A Ternary Read-only Memory Accelerator for LLM-powered Edge Intelligence
von: Guan, Hongyi, et al.
Veröffentlicht: (2026)
von: Guan, Hongyi, et al.
Veröffentlicht: (2026)
RISCBench: Benchmarking RISC-V Orchestration Efficiency in FPGA and FPGA-Like Computing Engines
von: Ojika, Dave, et al.
Veröffentlicht: (2025)
von: Ojika, Dave, et al.
Veröffentlicht: (2025)
RV-IM100: Quantifying ISA Extension, Datapath Width, and Pipeline Depth Trade-offs in RISC-V Microarchitectures
von: Kang, Hyunwoo
Veröffentlicht: (2026)
von: Kang, Hyunwoo
Veröffentlicht: (2026)
DSPE: An Energy-Efficient Edge Processor for DeepSeek Inference with MerkleTree-based Incremental Pruning, Multi-Stage Boothing Lookup and Dynamic Adaptive Posit Processing
von: Zhang, Yuhan, et al.
Veröffentlicht: (2026)
von: Zhang, Yuhan, et al.
Veröffentlicht: (2026)
basic_RV32s: An Open-Source Microarchitectural Roadmap for RISC-V RV32I
von: Kang, Hyun Woo, et al.
Veröffentlicht: (2025)
von: Kang, Hyun Woo, et al.
Veröffentlicht: (2025)
NeuroPDE: A Neuromorphic PDE Solver Based on Spintronic and Ferroelectric Devices
von: Fu, Siqing, et al.
Veröffentlicht: (2025)
von: Fu, Siqing, et al.
Veröffentlicht: (2025)
RISC-V Based TinyML Accelerator for Depthwise Separable Convolutions in Edge AI
von: Yildirim, Muhammed, et al.
Veröffentlicht: (2025)
von: Yildirim, Muhammed, et al.
Veröffentlicht: (2025)
ECOLogic: Enabling Circular, Obfuscated, and Adaptive Logic via eFPGA-Augmented SoCs
von: Tashdid, Ishraq, et al.
Veröffentlicht: (2025)
von: Tashdid, Ishraq, et al.
Veröffentlicht: (2025)
TokenStack: A Heterogeneous HBM-PIM Architecture and Runtime for Efficient LLM Inference
von: Li, Zhuoran, et al.
Veröffentlicht: (2026)
von: Li, Zhuoran, et al.
Veröffentlicht: (2026)
Biological Intuition on Digital Hardware: An RTL Implementation of Poisson-Encoded SNNs for Static Image Classification
von: Das, Debabrata, et al.
Veröffentlicht: (2026)
von: Das, Debabrata, et al.
Veröffentlicht: (2026)
A flexible framework for early power and timing comparison of time-multiplexed CGRA kernel executions
von: Aspros, Maxime Henri, et al.
Veröffentlicht: (2025)
von: Aspros, Maxime Henri, et al.
Veröffentlicht: (2025)
Time Domain Near Memory Computing Engine
von: Antal, Sarthak, et al.
Veröffentlicht: (2026)
von: Antal, Sarthak, et al.
Veröffentlicht: (2026)
Nonvolatile Charge-Domain Attention with HZO Ferroelectric Capacitors: A Simulation-Based Device-to-System Evaluation
von: Abouagour, Faris
Veröffentlicht: (2026)
von: Abouagour, Faris
Veröffentlicht: (2026)
InterPUF: Distributed Authentication via Physically Unclonable Functions and Multi-party Computation for Reconfigurable Interposers
von: Tashdid, Ishraq, et al.
Veröffentlicht: (2026)
von: Tashdid, Ishraq, et al.
Veröffentlicht: (2026)
Ten-Four: An Open-Source Fused Dot Product Unit for Mixed-Precision GPGPU Tensor Cores
von: Rout, Nikhil, et al.
Veröffentlicht: (2025)
von: Rout, Nikhil, et al.
Veröffentlicht: (2025)
Factor Machine: Mixed-signal Architecture for Fine-Grained Graph-Based Computing
von: Dudek, Piotr
Veröffentlicht: (2024)
von: Dudek, Piotr
Veröffentlicht: (2024)
Lean Attention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers
von: Sanovar, Rya, et al.
Veröffentlicht: (2024)
von: Sanovar, Rya, et al.
Veröffentlicht: (2024)
GainSight: A Unified Framework for Data Lifetime Profiling and Heterogeneous Memory Composition
von: Li, Peijing, et al.
Veröffentlicht: (2025)
von: Li, Peijing, et al.
Veröffentlicht: (2025)
CGRA4ML: A Hardware/Software Framework to Implement Neural Networks for Scientific Edge Computing
von: Abarajithan, G, et al.
Veröffentlicht: (2024)
von: Abarajithan, G, et al.
Veröffentlicht: (2024)
Minimal Neuron Circuits -- Part I: Resonators
von: Nabil, Amr, et al.
Veröffentlicht: (2025)
von: Nabil, Amr, et al.
Veröffentlicht: (2025)
Minimal Neuron Circuits: Bursters
von: Nabil, Amr, et al.
Veröffentlicht: (2025)
von: Nabil, Amr, et al.
Veröffentlicht: (2025)
FlexiBit: Fully Flexible Precision Bit-parallel Accelerator Architecture for Arbitrary Mixed Precision AI
von: Tahmasebi, Faraz, et al.
Veröffentlicht: (2024)
von: Tahmasebi, Faraz, et al.
Veröffentlicht: (2024)
D-com: Accelerating Iterative Processing to Enable Low-rank Decomposition of Activations
von: Tahmasebi, Faraz, et al.
Veröffentlicht: (2025)
von: Tahmasebi, Faraz, et al.
Veröffentlicht: (2025)
Not All Thoughts Need HBM: Semantics-Aware Memory Hierarchy for LLM Reasoning
von: Yuan, Aojie, et al.
Veröffentlicht: (2026)
von: Yuan, Aojie, et al.
Veröffentlicht: (2026)
An Integrated UVM-TLM Co-Simulation Framework for RISC-V Functional Verification and Performance Evaluation
von: Qiu, Ruizhi, et al.
Veröffentlicht: (2025)
von: Qiu, Ruizhi, et al.
Veröffentlicht: (2025)
Lincoln AI Computing Survey (LAICS) and Trends
von: Reuther, Albert, et al.
Veröffentlicht: (2025)
von: Reuther, Albert, et al.
Veröffentlicht: (2025)
A 65 nm Bayesian Neural Network Accelerator with 360 fJ/Sample In-Word GRNG for AI Uncertainty Estimation
von: Enciso, Zephan M., et al.
Veröffentlicht: (2025)
von: Enciso, Zephan M., et al.
Veröffentlicht: (2025)
A Per-Access Upper Bound for Shared-Resource Interference in Direct-Mapped Multicore Architectures
von: Pedroni, Felipe T.
Veröffentlicht: (2026)
von: Pedroni, Felipe T.
Veröffentlicht: (2026)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
von: Jo, Myeong Jun
Veröffentlicht: (2026)
von: Jo, Myeong Jun
Veröffentlicht: (2026)
Reed-Muller Error-Correction Code Encoder for SFQ-to-CMOS Interface Circuits
von: Mustafa, Yerzhan, et al.
Veröffentlicht: (2026)
von: Mustafa, Yerzhan, et al.
Veröffentlicht: (2026)
FPGA-Accelerated RISC-V ISA Extensions for Efficient Neural Network Inference on Edge Devices
von: Parameshwara, Arya, et al.
Veröffentlicht: (2025)
von: Parameshwara, Arya, et al.
Veröffentlicht: (2025)
Accelerating Precise End-to-End Simulation: Latency-Sensitive Many-core System Modeling
von: Li, Yinrong, et al.
Veröffentlicht: (2026)
von: Li, Yinrong, et al.
Veröffentlicht: (2026)
RayFlex: An Open-Source RTL Implementation of the Hardware Ray Tracer Datapath
von: Shen, Fangjia, et al.
Veröffentlicht: (2024)
von: Shen, Fangjia, et al.
Veröffentlicht: (2024)
RedMulE-FT: A Reconfigurable Fault-Tolerant Matrix Multiplication Engine
von: Wiese, Philip, et al.
Veröffentlicht: (2025)
von: Wiese, Philip, et al.
Veröffentlicht: (2025)
Toward a Universal GPU Instruction Set Architecture: A Cross-Vendor Analysis of Hardware-Invariant Computational Primitives in Parallel Processors
von: Abraham, Ojima, et al.
Veröffentlicht: (2026)
von: Abraham, Ojima, et al.
Veröffentlicht: (2026)
A Comparative Analysis of ARM and x86-64 Laptop-Class Processors: Architecture, Assembly-Level Performance, and Energy Efficiency
von: Özyılmaz, Mustafa Mert
Veröffentlicht: (2026)
von: Özyılmaz, Mustafa Mert
Veröffentlicht: (2026)
Lightweight Error-Correction Code Encoders in Superconducting Electronic Systems
von: Mustafa, Yerzhan, et al.
Veröffentlicht: (2025)
von: Mustafa, Yerzhan, et al.
Veröffentlicht: (2025)
SynapticCore-X: A Modular Neural Processing Architecture for Low-Cost FPGA Acceleration
von: Parameshwara, Arya
Veröffentlicht: (2025)
von: Parameshwara, Arya
Veröffentlicht: (2025)
Wattchmen: Watching the Wattchers -- High Fidelity, Flexible GPU Energy Modeling
von: Tran, Brandon, et al.
Veröffentlicht: (2026)
von: Tran, Brandon, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Design and Implementation of an FPGA-Based Hardware Accelerator for Transformer
von: Li, Richie, et al.
Veröffentlicht: (2025) -
TOM: A Ternary Read-only Memory Accelerator for LLM-powered Edge Intelligence
von: Guan, Hongyi, et al.
Veröffentlicht: (2026) -
RISCBench: Benchmarking RISC-V Orchestration Efficiency in FPGA and FPGA-Like Computing Engines
von: Ojika, Dave, et al.
Veröffentlicht: (2025) -
RV-IM100: Quantifying ISA Extension, Datapath Width, and Pipeline Depth Trade-offs in RISC-V Microarchitectures
von: Kang, Hyunwoo
Veröffentlicht: (2026) -
DSPE: An Energy-Efficient Edge Processor for DeepSeek Inference with MerkleTree-based Incremental Pruning, Multi-Stage Boothing Lookup and Dynamic Adaptive Posit Processing
von: Zhang, Yuhan, et al.
Veröffentlicht: (2026)