LlamaF: An Efficient Llama2 Architecture Accelerator on Embedded FPGAs
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Han, Li, Yutong, Ji, Shihao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Energy Efficient LSTM Accelerators for Embedded FPGAs through Parameterised Architecture Design
by: Qian, Chao, et al.
Published: (2026)
by: Qian, Chao, et al.
Published: (2026)
HLSTransform: Energy-Efficient Llama 2 Inference on FPGAs Via High Level Synthesis
by: He, Andy, et al.
Published: (2024)
by: He, Andy, et al.
Published: (2024)
Accelerating Boolean Constraint Propagation for Efficient SAT-Solving on FPGAs
by: Govindasamy, Hariprasadh, et al.
Published: (2024)
by: Govindasamy, Hariprasadh, et al.
Published: (2024)
Efficient Approaches for GEMM Acceleration on Leading AI-Optimized FPGAs
by: Taka, Endri, et al.
Published: (2024)
by: Taka, Endri, et al.
Published: (2024)
Duet: Creating Harmony between Processors and Embedded FPGAs
by: Li, Ang, et al.
Published: (2023)
by: Li, Ang, et al.
Published: (2023)
An Energy-Efficient Artefact Detection Accelerator on FPGAs for Hyper-Spectral Satellite Imagery
by: Castelino, Cornell, et al.
Published: (2024)
by: Castelino, Cornell, et al.
Published: (2024)
Memory-efficient Sketch Acceleration for Handling Large Network Flows on FPGAs
by: Han, Zhaoyang, et al.
Published: (2025)
by: Han, Zhaoyang, et al.
Published: (2025)
Leveraging Application-Specific Knowledge for Energy-Efficient Deep Learning Accelerators on Resource-Constrained FPGAs
by: Qian, Chao
Published: (2025)
by: Qian, Chao
Published: (2025)
SWAT: Scalable and Efficient Window Attention-based Transformers Acceleration on FPGAs
by: Bai, Zhenyu, et al.
Published: (2024)
by: Bai, Zhenyu, et al.
Published: (2024)
Theoretical Analysis of the Efficient-Memory Matrix Storage Method for Quantum Emulation Accelerators with Gate Fusion on FPGAs
by: Le, Tran Xuan Hieu, et al.
Published: (2024)
by: Le, Tran Xuan Hieu, et al.
Published: (2024)
TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs
by: Qiao, Ye, et al.
Published: (2025)
by: Qiao, Ye, et al.
Published: (2025)
Hardware-Efficient Accurate 4-bit Multiplier for Xilinx 7 Series FPGAs
by: Kida, Misaki, et al.
Published: (2025)
by: Kida, Misaki, et al.
Published: (2025)
Efficient Message Passing Architecture for GCN Training on HBM-based FPGAs with Orthogonal Topology On-Chip Networks
by: Wu, Qizhe, et al.
Published: (2024)
by: Wu, Qizhe, et al.
Published: (2024)
Coyote v2: Raising the Level of Abstraction for Data Center FPGAs
by: Ramhorst, Benjamin, et al.
Published: (2025)
by: Ramhorst, Benjamin, et al.
Published: (2025)
DORA: Dataflow-Instruction Orchestration Architecture for DNN Acceleration
by: Chen, Xingzhen, et al.
Published: (2026)
by: Chen, Xingzhen, et al.
Published: (2026)
Wavelet Based Frequency Detection Using FPGAs
by: Hill, Caleb, et al.
Published: (2024)
by: Hill, Caleb, et al.
Published: (2024)
Dynamic Tsetlin Machine Accelerators for On-Chip Training at the Edge using FPGAs
by: Mao, Gang, et al.
Published: (2025)
by: Mao, Gang, et al.
Published: (2025)
Studying the Degradation of Propagation Delay on FPGAs at the European XFEL
by: Lanzieri, Leandro, et al.
Published: (2024)
by: Lanzieri, Leandro, et al.
Published: (2024)
Interconnect-Aware Logic Resynthesis for Multi-Die FPGAs
by: Wang, Xiaoke, et al.
Published: (2026)
by: Wang, Xiaoke, et al.
Published: (2026)
Design and Implementation of Washing Machine HUD Using FPGAs
by: Stites, Norman, et al.
Published: (2025)
by: Stites, Norman, et al.
Published: (2025)
Optimized Spatial Architecture Mapping Flow for Transformer Accelerators
by: Xu, Haocheng, et al.
Published: (2024)
by: Xu, Haocheng, et al.
Published: (2024)
TeLLMe v2: An Efficient End-to-End Ternary LLM Prefill and Decode Accelerator with Table-Lookup Matmul on Edge FPGAs
by: Qiao, Ye, et al.
Published: (2025)
by: Qiao, Ye, et al.
Published: (2025)
FILCO: Flexible Composing Architecture with Real-Time Reconfigurability for DNN Acceleration
by: Chen, Xingzhen, et al.
Published: (2026)
by: Chen, Xingzhen, et al.
Published: (2026)
Fast Generation of Custom Floating-Point Spatial Filters on FPGAs
by: Campos, Nelson, et al.
Published: (2024)
by: Campos, Nelson, et al.
Published: (2024)
Small Logic-based Multipliers with Incomplete Sub-Multipliers for FPGAs
by: Böttcher, Andreas, et al.
Published: (2024)
by: Böttcher, Andreas, et al.
Published: (2024)
RidgeWalker: Perfectly Pipelined Graph Random Walks on FPGAs
by: Tan, Hongshi, et al.
Published: (2026)
by: Tan, Hongshi, et al.
Published: (2026)
smallNet: Implementation of a convolutional layer in tiny FPGAs
by: Bascuñán, Fernanda Zapata, et al.
Published: (2025)
by: Bascuñán, Fernanda Zapata, et al.
Published: (2025)
GreenFPGA: Evaluating FPGAs as Environmentally Sustainable Computing Solutions
by: Sudarshan, Chetan Choppali, et al.
Published: (2023)
by: Sudarshan, Chetan Choppali, et al.
Published: (2023)
Lightweight Congruence Profiling for Early Design Exploration of Heterogeneous FPGAs
by: Boston, Allen, et al.
Published: (2025)
by: Boston, Allen, et al.
Published: (2025)
DPUConfig: Optimizing ML Inference in FPGAs Using Reinforcement Learning
by: Patras, Alexandros, et al.
Published: (2026)
by: Patras, Alexandros, et al.
Published: (2026)
Transitive Array: An Efficient GEMM Accelerator with Result Reuse
by: Guo, Cong, et al.
Published: (2025)
by: Guo, Cong, et al.
Published: (2025)
Critical Path Aware Timing-Driven Global Placement for Large-Scale Heterogeneous FPGAs
by: Jiang, He, et al.
Published: (2025)
by: Jiang, He, et al.
Published: (2025)
ONNX-to-Hardware Design Flow for Adaptive Neural-Network Inference on FPGAs
by: Manca, Federico, et al.
Published: (2024)
by: Manca, Federico, et al.
Published: (2024)
A Survey on LUT-based Deep Neural Networks Implemented in FPGAs
by: Guo, Zeyu
Published: (2025)
by: Guo, Zeyu
Published: (2025)
Escaping Flatland: A Placement Flow for Enabling 3D FPGAs
by: Hao, Cong, et al.
Published: (2026)
by: Hao, Cong, et al.
Published: (2026)
A Quarter of a Century of Neuromorphic Architectures on FPGAs -- an Overview
by: Szczerek, Wiktor J., et al.
Published: (2025)
by: Szczerek, Wiktor J., et al.
Published: (2025)
Hummingbird: A Smaller and Faster Large Language Model Accelerator on Embedded FPGA
by: Li, Jindong, et al.
Published: (2025)
by: Li, Jindong, et al.
Published: (2025)
Analyzing the Single Event Upset Vulnerability of Binarized Neural Networks on SRAM FPGAs
by: Souvatzoglou, Ioanna, et al.
Published: (2024)
by: Souvatzoglou, Ioanna, et al.
Published: (2024)
A Resource-Driven Approach for Implementing CNNs on FPGAs Using Adaptive IPs
by: Magalhães, Philippe, et al.
Published: (2025)
by: Magalhães, Philippe, et al.
Published: (2025)
EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture
by: Duan, Bowen, et al.
Published: (2026)
by: Duan, Bowen, et al.
Published: (2026)
Similar Items
-
Energy Efficient LSTM Accelerators for Embedded FPGAs through Parameterised Architecture Design
by: Qian, Chao, et al.
Published: (2026) -
HLSTransform: Energy-Efficient Llama 2 Inference on FPGAs Via High Level Synthesis
by: He, Andy, et al.
Published: (2024) -
Accelerating Boolean Constraint Propagation for Efficient SAT-Solving on FPGAs
by: Govindasamy, Hariprasadh, et al.
Published: (2024) -
Efficient Approaches for GEMM Acceleration on Leading AI-Optimized FPGAs
by: Taka, Endri, et al.
Published: (2024) -
Duet: Creating Harmony between Processors and Embedded FPGAs
by: Li, Ang, et al.
Published: (2023)