Hummingbird: A Smaller and Faster Large Language Model Accelerator on Embedded FPGA
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Jindong, Li, Tenglong, Chen, Ruiqi, Shen, Guobin, Zhao, Dongcheng, Zhang, Qian, Zeng, Yi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Pushing up to the Limit of Memory Bandwidth and Capacity Utilization for Efficient LLM Decoding on Embedded FPGA
von: Li, Jindong, et al.
Veröffentlicht: (2025)
von: Li, Jindong, et al.
Veröffentlicht: (2025)
FireFly-P: FPGA-Accelerated Spiking Neural Network Plasticity for Robust Adaptive Control
von: Li, Tenglong, et al.
Veröffentlicht: (2026)
von: Li, Tenglong, et al.
Veröffentlicht: (2026)
Revealing Untapped DSP Optimization Potentials for FPGA-Based Systolic Matrix Engines
von: Li, Jindong, et al.
Veröffentlicht: (2024)
von: Li, Jindong, et al.
Veröffentlicht: (2024)
FireFly-S: Exploiting Dual-Side Sparsity for Spiking Neural Networks Acceleration with Reconfigurable Spatial Architecture
von: Li, Tenglong, et al.
Veröffentlicht: (2024)
von: Li, Tenglong, et al.
Veröffentlicht: (2024)
FireFly-T: High-Throughput Sparsity Exploitation for Spiking Transformer Acceleration with Dual-Engine Overlay Architecture
von: Li, Tenglong, et al.
Veröffentlicht: (2025)
von: Li, Tenglong, et al.
Veröffentlicht: (2025)
EdgeLLM: A Highly Efficient CPU-FPGA Heterogeneous Edge Accelerator for Large Language Models
von: Huang, Mingqiang, et al.
Veröffentlicht: (2024)
von: Huang, Mingqiang, et al.
Veröffentlicht: (2024)
A High-Throughput FPGA Accelerator for Lightweight CNNs With Balanced Dataflow
von: Zhao, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Zhao, Zhiyuan, et al.
Veröffentlicht: (2024)
SpeedLLM: An FPGA Co-design of Large Language Model Inference Accelerator
von: Wang, Peipei, et al.
Veröffentlicht: (2025)
von: Wang, Peipei, et al.
Veröffentlicht: (2025)
Graphitron: A Domain Specific Language for FPGA-based Graph Processing Accelerator Generation
von: Zhang, Xinmiao, et al.
Veröffentlicht: (2024)
von: Zhang, Xinmiao, et al.
Veröffentlicht: (2024)
SkipOPU: An FPGA-based Overlay Processor for Large Language Models with Dynamically Allocated Computation
von: He, Zicheng, et al.
Veröffentlicht: (2026)
von: He, Zicheng, et al.
Veröffentlicht: (2026)
Managing Hybrid Solid-State Drives Using Large Language Models
von: Wei, Qian, et al.
Veröffentlicht: (2025)
von: Wei, Qian, et al.
Veröffentlicht: (2025)
An Optimizing Framework on MLIR for Efficient FPGA-based Accelerator Generation
von: Zhang, Weichuang, et al.
Veröffentlicht: (2024)
von: Zhang, Weichuang, et al.
Veröffentlicht: (2024)
HAAN: A Holistic Approach for Accelerating Normalization Operations in Large Language Models
von: Peng, Tianfan, et al.
Veröffentlicht: (2025)
von: Peng, Tianfan, et al.
Veröffentlicht: (2025)
Holistic Optimization Framework for FPGA Accelerators
von: Pouget, Stéphane, et al.
Veröffentlicht: (2025)
von: Pouget, Stéphane, et al.
Veröffentlicht: (2025)
SpecMamba: Accelerating Mamba Inference on FPGA with Speculative Decoding
von: Zhong, Linfeng, et al.
Veröffentlicht: (2025)
von: Zhong, Linfeng, et al.
Veröffentlicht: (2025)
FPGA-Optimized Hardware Accelerator for Fast Fourier Transform and Singular Value Decomposition in AI
von: Ding, Hong, et al.
Veröffentlicht: (2025)
von: Ding, Hong, et al.
Veröffentlicht: (2025)
TurboFuzz: FPGA Accelerated Hardware Fuzzing for Processor Agile Verification
von: Zhong, Yang, et al.
Veröffentlicht: (2025)
von: Zhong, Yang, et al.
Veröffentlicht: (2025)
SimulatorCoder: DNN Accelerator Simulator Code Generation and Optimization via Large Language Models
von: Xia, Yuhuan, et al.
Veröffentlicht: (2026)
von: Xia, Yuhuan, et al.
Veröffentlicht: (2026)
A Reconfigurable Framework for AI-FPGA Agent Integration and Acceleration
von: Yunusoglu, Aybars, et al.
Veröffentlicht: (2026)
von: Yunusoglu, Aybars, et al.
Veröffentlicht: (2026)
Bombyx: OpenCilk Compilation for FPGA Hardware Acceleration
von: Shahawy, Mohamed, et al.
Veröffentlicht: (2025)
von: Shahawy, Mohamed, et al.
Veröffentlicht: (2025)
Implementation and Analysis of Thermometer Encoding in DWN FPGA Accelerators
von: Mecik, Michael, et al.
Veröffentlicht: (2025)
von: Mecik, Michael, et al.
Veröffentlicht: (2025)
Design and Implementation of BNN-Based Object Detection on FPGA
von: Zhao, Xuyu, et al.
Veröffentlicht: (2026)
von: Zhao, Xuyu, et al.
Veröffentlicht: (2026)
Swift: A Multi-FPGA Framework for Scaling Up Accelerated Graph Analytics
von: Jaiyeoba, Oluwole, et al.
Veröffentlicht: (2024)
von: Jaiyeoba, Oluwole, et al.
Veröffentlicht: (2024)
SuperUROP: An FPGA-Based Spatial Accelerator for Sparse Matrix Operations
von: Parthasarathy, Rishab
Veröffentlicht: (2025)
von: Parthasarathy, Rishab
Veröffentlicht: (2025)
An Irredundant and Compressed Data Layout to Optimize Bandwidth Utilization of FPGA Accelerators
von: Ferry, Corentin, et al.
Veröffentlicht: (2024)
von: Ferry, Corentin, et al.
Veröffentlicht: (2024)
Embedded FPGA Acceleration of Brain-Like Neural Networks: Online Learning to Scalable Inference
von: Hafiz, Muhammad Ihsan Al, et al.
Veröffentlicht: (2025)
von: Hafiz, Muhammad Ihsan Al, et al.
Veröffentlicht: (2025)
Energy Efficient LSTM Accelerators for Embedded FPGAs through Parameterised Architecture Design
von: Qian, Chao, et al.
Veröffentlicht: (2026)
von: Qian, Chao, et al.
Veröffentlicht: (2026)
FASE: FPGA-Assisted Syscall Emulation for Rapid End-to-End Processor Performance Validation
von: Meng, Chengzhen, et al.
Veröffentlicht: (2025)
von: Meng, Chengzhen, et al.
Veröffentlicht: (2025)
FAST-Prefill: FPGA Accelerated Sparse Attention for Long Context LLM Prefill
von: Jayanth, Rakshith, et al.
Veröffentlicht: (2026)
von: Jayanth, Rakshith, et al.
Veröffentlicht: (2026)
Facial Expression Recognition System Using DNN Accelerator with Multi-threading on FPGA
von: Ando, Takuto, et al.
Veröffentlicht: (2025)
von: Ando, Takuto, et al.
Veröffentlicht: (2025)
always_comm: An FPGA-based Hardware Accelerator for Audio/Video Compression and Transmission
von: Parthasarathy, Rishab, et al.
Veröffentlicht: (2025)
von: Parthasarathy, Rishab, et al.
Veröffentlicht: (2025)
Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inference
von: Chen, Hongzheng, et al.
Veröffentlicht: (2023)
von: Chen, Hongzheng, et al.
Veröffentlicht: (2023)
HPIM: Heterogeneous Processing-In-Memory-based Accelerator for Large Language Models Inference
von: Duan, Cenlin, et al.
Veröffentlicht: (2025)
von: Duan, Cenlin, et al.
Veröffentlicht: (2025)
ZynqParrot: A Scale-Down Approach to Cycle-Accurate, FPGA-Accelerated Co-Emulation
von: Ruelas-Petrisko, Daniel, et al.
Veröffentlicht: (2025)
von: Ruelas-Petrisko, Daniel, et al.
Veröffentlicht: (2025)
FPGA Acceleration of Image Reconstruction for Real-Time Photoacoustic Tomography
von: Gao, Zijian, et al.
Veröffentlicht: (2022)
von: Gao, Zijian, et al.
Veröffentlicht: (2022)
Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective
von: Li, Jinhao, et al.
Veröffentlicht: (2024)
von: Li, Jinhao, et al.
Veröffentlicht: (2024)
LlamaF: An Efficient Llama2 Architecture Accelerator on Embedded FPGAs
von: Xu, Han, et al.
Veröffentlicht: (2024)
von: Xu, Han, et al.
Veröffentlicht: (2024)
Late Breaking Result: FPGA-Based Emulation and Fault Injection for CNN Inference Accelerators
von: Masar, Filip, et al.
Veröffentlicht: (2025)
von: Masar, Filip, et al.
Veröffentlicht: (2025)
SIRA: Scaled-Integer Range Analysis for Optimizing FPGA Dataflow Neural Network Accelerators
von: Umuroglu, Yaman, et al.
Veröffentlicht: (2025)
von: Umuroglu, Yaman, et al.
Veröffentlicht: (2025)
Systolic Sparse Tensor Slices: FPGA Building Blocks for Sparse and Dense AI Acceleration
von: Taka, Endri, et al.
Veröffentlicht: (2025)
von: Taka, Endri, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Pushing up to the Limit of Memory Bandwidth and Capacity Utilization for Efficient LLM Decoding on Embedded FPGA
von: Li, Jindong, et al.
Veröffentlicht: (2025) -
FireFly-P: FPGA-Accelerated Spiking Neural Network Plasticity for Robust Adaptive Control
von: Li, Tenglong, et al.
Veröffentlicht: (2026) -
Revealing Untapped DSP Optimization Potentials for FPGA-Based Systolic Matrix Engines
von: Li, Jindong, et al.
Veröffentlicht: (2024) -
FireFly-S: Exploiting Dual-Side Sparsity for Spiking Neural Networks Acceleration with Reconfigurable Spatial Architecture
von: Li, Tenglong, et al.
Veröffentlicht: (2024) -
FireFly-T: High-Throughput Sparsity Exploitation for Spiking Transformer Acceleration with Dual-Engine Overlay Architecture
von: Li, Tenglong, et al.
Veröffentlicht: (2025)