FlightLLM: Efficient Large Language Model Inference with a Complete Mapping Flow on FPGAs
Fuente:
arXiv
Saved in:
| Main Authors: | Zeng, Shulin, Liu, Jun, Dai, Guohao, Yang, Xinhao, Fu, Tianyu, Wang, Hongyi, Ma, Wenheng, Sun, Hanbo, Li, Shiyao, Huang, Zixiao, Dai, Yadong, Li, Jintao, Wang, Zehao, Zhang, Ruoyu, Wen, Kairui, Ning, Xuefei, Wang, Yu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Duet: Creating Harmony between Processors and Embedded FPGAs
by: Li, Ang, et al.
Published: (2023)
by: Li, Ang, et al.
Published: (2023)
Critical Path Aware Timing-Driven Global Placement for Large-Scale Heterogeneous FPGAs
by: Jiang, He, et al.
Published: (2025)
by: Jiang, He, et al.
Published: (2025)
DPUConfig: Optimizing ML Inference in FPGAs Using Reinforcement Learning
by: Patras, Alexandros, et al.
Published: (2026)
by: Patras, Alexandros, et al.
Published: (2026)
Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective
by: Li, Jinhao, et al.
Published: (2024)
by: Li, Jinhao, et al.
Published: (2024)
ONNX-to-Hardware Design Flow for Adaptive Neural-Network Inference on FPGAs
by: Manca, Federico, et al.
Published: (2024)
by: Manca, Federico, et al.
Published: (2024)
LUT-LLM: Efficient Large Language Model Inference with Memory-based Computations on FPGAs
by: He, Zifan, et al.
Published: (2025)
by: He, Zifan, et al.
Published: (2025)
Interconnect-Aware Logic Resynthesis for Multi-Die FPGAs
by: Wang, Xiaoke, et al.
Published: (2026)
by: Wang, Xiaoke, et al.
Published: (2026)
PolyThrottle: Energy-efficient Neural Network Inference on Edge Devices
by: Yan, Minghao, et al.
Published: (2023)
by: Yan, Minghao, et al.
Published: (2023)
Data-Rate-Aware High-Speed CNN Inference on FPGAs
by: Habermann, Tobias, et al.
Published: (2026)
by: Habermann, Tobias, et al.
Published: (2026)
LlamaF: An Efficient Llama2 Architecture Accelerator on Embedded FPGAs
by: Xu, Han, et al.
Published: (2024)
by: Xu, Han, et al.
Published: (2024)
Testing and Fault Tolerance Techniques for CNT-Based FPGAs
by: Lu, Siyuan, et al.
Published: (2025)
by: Lu, Siyuan, et al.
Published: (2025)
MARCA: Mamba Accelerator with ReConfigurable Architecture
by: Li, Jinhao, et al.
Published: (2024)
by: Li, Jinhao, et al.
Published: (2024)
Enabling Long FFT Convolutions on Memory-Constrained FPGAs via Chunking
by: Wang, Peter, et al.
Published: (2025)
by: Wang, Peter, et al.
Published: (2025)
QUADOL: A Quality-Driven Approximate Logic Synthesis Method Exploiting Dual-Output LUTs for Modern FPGAs
by: Shi, Jian, et al.
Published: (2024)
by: Shi, Jian, et al.
Published: (2024)
SoftmAP: Software-Hardware Co-design for Integer-Only Softmax on Associative Processors
by: Rakka, Mariam, et al.
Published: (2024)
by: Rakka, Mariam, et al.
Published: (2024)
Wavelet Based Frequency Detection Using FPGAs
by: Hill, Caleb, et al.
Published: (2024)
by: Hill, Caleb, et al.
Published: (2024)
H2PIPE: High throughput CNN Inference on FPGAs with High-Bandwidth Memory
by: Doumet, Mario, et al.
Published: (2024)
by: Doumet, Mario, et al.
Published: (2024)
PD-Swap: Prefill-Decode Logic Swapping for End-to-End LLM Inference on Edge FPGAs via Dynamic Partial Reconfiguration
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Runtime Tunable Tsetlin Machines for Edge Inference on eFPGAs
by: Rahman, Tousif, et al.
Published: (2025)
by: Rahman, Tousif, et al.
Published: (2025)
Studying the Degradation of Propagation Delay on FPGAs at the European XFEL
by: Lanzieri, Leandro, et al.
Published: (2024)
by: Lanzieri, Leandro, et al.
Published: (2024)
Design and Implementation of Washing Machine HUD Using FPGAs
by: Stites, Norman, et al.
Published: (2025)
by: Stites, Norman, et al.
Published: (2025)
RidgeWalker: Perfectly Pipelined Graph Random Walks on FPGAs
by: Tan, Hongshi, et al.
Published: (2026)
by: Tan, Hongshi, et al.
Published: (2026)
Fast Generation of Custom Floating-Point Spatial Filters on FPGAs
by: Campos, Nelson, et al.
Published: (2024)
by: Campos, Nelson, et al.
Published: (2024)
smallNet: Implementation of a convolutional layer in tiny FPGAs
by: Bascuñán, Fernanda Zapata, et al.
Published: (2025)
by: Bascuñán, Fernanda Zapata, et al.
Published: (2025)
GreenFPGA: Evaluating FPGAs as Environmentally Sustainable Computing Solutions
by: Sudarshan, Chetan Choppali, et al.
Published: (2023)
by: Sudarshan, Chetan Choppali, et al.
Published: (2023)
Small Logic-based Multipliers with Incomplete Sub-Multipliers for FPGAs
by: Böttcher, Andreas, et al.
Published: (2024)
by: Böttcher, Andreas, et al.
Published: (2024)
Lightweight Congruence Profiling for Early Design Exploration of Heterogeneous FPGAs
by: Boston, Allen, et al.
Published: (2025)
by: Boston, Allen, et al.
Published: (2025)
Accelerating Boolean Constraint Propagation for Efficient SAT-Solving on FPGAs
by: Govindasamy, Hariprasadh, et al.
Published: (2024)
by: Govindasamy, Hariprasadh, et al.
Published: (2024)
Efficient Approaches for GEMM Acceleration on Leading AI-Optimized FPGAs
by: Taka, Endri, et al.
Published: (2024)
by: Taka, Endri, et al.
Published: (2024)
GraphMatch: Subgraph Query Processing on FPGAs
by: Dann, Jonas, et al.
Published: (2024)
by: Dann, Jonas, et al.
Published: (2024)
SWAT: Scalable and Efficient Window Attention-based Transformers Acceleration on FPGAs
by: Bai, Zhenyu, et al.
Published: (2024)
by: Bai, Zhenyu, et al.
Published: (2024)
SHIELD: A Segmented Hierarchical Memory Architecture for Energy-Efficient LLM Inference on Edge NPUs
by: Zhang, Jintao, et al.
Published: (2026)
by: Zhang, Jintao, et al.
Published: (2026)
TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs
by: Qiao, Ye, et al.
Published: (2025)
by: Qiao, Ye, et al.
Published: (2025)
A Survey on LUT-based Deep Neural Networks Implemented in FPGAs
by: Guo, Zeyu
Published: (2025)
by: Guo, Zeyu
Published: (2025)
Coyote v2: Raising the Level of Abstraction for Data Center FPGAs
by: Ramhorst, Benjamin, et al.
Published: (2025)
by: Ramhorst, Benjamin, et al.
Published: (2025)
Escaping Flatland: A Placement Flow for Enabling 3D FPGAs
by: Hao, Cong, et al.
Published: (2026)
by: Hao, Cong, et al.
Published: (2026)
Preemption-Enhanced Benchmark Suite for FPGAs
by: Malik, Arsalan Ali, et al.
Published: (2025)
by: Malik, Arsalan Ali, et al.
Published: (2025)
A Resource-Driven Approach for Implementing CNNs on FPGAs Using Adaptive IPs
by: Magalhães, Philippe, et al.
Published: (2025)
by: Magalhães, Philippe, et al.
Published: (2025)
Hardware-Efficient Accurate 4-bit Multiplier for Xilinx 7 Series FPGAs
by: Kida, Misaki, et al.
Published: (2025)
by: Kida, Misaki, et al.
Published: (2025)
Energy Efficient LSTM Accelerators for Embedded FPGAs through Parameterised Architecture Design
by: Qian, Chao, et al.
Published: (2026)
by: Qian, Chao, et al.
Published: (2026)
Similar Items
-
Duet: Creating Harmony between Processors and Embedded FPGAs
by: Li, Ang, et al.
Published: (2023) -
Critical Path Aware Timing-Driven Global Placement for Large-Scale Heterogeneous FPGAs
by: Jiang, He, et al.
Published: (2025) -
DPUConfig: Optimizing ML Inference in FPGAs Using Reinforcement Learning
by: Patras, Alexandros, et al.
Published: (2026) -
Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective
by: Li, Jinhao, et al.
Published: (2024) -
ONNX-to-Hardware Design Flow for Adaptive Neural-Network Inference on FPGAs
by: Manca, Federico, et al.
Published: (2024)