HLSTransform: Energy-Efficient Llama 2 Inference on FPGAs Via High Level Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | He, Andy, Key, Darren, Bulling, Mason, Chang, Andrew, Shapiro, Skyler, Lee, Everett |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LlamaF: An Efficient Llama2 Architecture Accelerator on Embedded FPGAs
by: Xu, Han, et al.
Published: (2024)
by: Xu, Han, et al.
Published: (2024)
LUT-LLM: Efficient Large Language Model Inference with Memory-based Computations on FPGAs
by: He, Zifan, et al.
Published: (2025)
by: He, Zifan, et al.
Published: (2025)
FlightLLM: Efficient Large Language Model Inference with a Complete Mapping Flow on FPGAs
by: Zeng, Shulin, et al.
Published: (2024)
by: Zeng, Shulin, et al.
Published: (2024)
SWAT: Scalable and Efficient Window Attention-based Transformers Acceleration on FPGAs
by: Bai, Zhenyu, et al.
Published: (2024)
by: Bai, Zhenyu, et al.
Published: (2024)
Energy Efficient LSTM Accelerators for Embedded FPGAs through Parameterised Architecture Design
by: Qian, Chao, et al.
Published: (2026)
by: Qian, Chao, et al.
Published: (2026)
An Energy-Efficient Artefact Detection Accelerator on FPGAs for Hyper-Spectral Satellite Imagery
by: Castelino, Cornell, et al.
Published: (2024)
by: Castelino, Cornell, et al.
Published: (2024)
Runtime Tunable Tsetlin Machines for Edge Inference on eFPGAs
by: Rahman, Tousif, et al.
Published: (2025)
by: Rahman, Tousif, et al.
Published: (2025)
Are LLMs Any Good for High-Level Synthesis?
by: Liao, Yuchao, et al.
Published: (2024)
by: Liao, Yuchao, et al.
Published: (2024)
DPUConfig: Optimizing ML Inference in FPGAs Using Reinforcement Learning
by: Patras, Alexandros, et al.
Published: (2026)
by: Patras, Alexandros, et al.
Published: (2026)
Leveraging Application-Specific Knowledge for Energy-Efficient Deep Learning Accelerators on Resource-Constrained FPGAs
by: Qian, Chao
Published: (2025)
by: Qian, Chao
Published: (2025)
Resource Utilization of Differentiable Logic Gate Networks Deployed on FPGAs
by: Wormald, Stephen, et al.
Published: (2026)
by: Wormald, Stephen, et al.
Published: (2026)
ONNX-to-Hardware Design Flow for Adaptive Neural-Network Inference on FPGAs
by: Manca, Federico, et al.
Published: (2024)
by: Manca, Federico, et al.
Published: (2024)
ELSA: An ELastic SNN Inference Architecture for Efficient Neuromorphic Computing
by: You, Kang, et al.
Published: (2026)
by: You, Kang, et al.
Published: (2026)
QUADOL: A Quality-Driven Approximate Logic Synthesis Method Exploiting Dual-Output LUTs for Modern FPGAs
by: Shi, Jian, et al.
Published: (2024)
by: Shi, Jian, et al.
Published: (2024)
Coyote v2: Raising the Level of Abstraction for Data Center FPGAs
by: Ramhorst, Benjamin, et al.
Published: (2025)
by: Ramhorst, Benjamin, et al.
Published: (2025)
Data-Rate-Aware High-Speed CNN Inference on FPGAs
by: Habermann, Tobias, et al.
Published: (2026)
by: Habermann, Tobias, et al.
Published: (2026)
Accelerating Boolean Constraint Propagation for Efficient SAT-Solving on FPGAs
by: Govindasamy, Hariprasadh, et al.
Published: (2024)
by: Govindasamy, Hariprasadh, et al.
Published: (2024)
Efficient Approaches for GEMM Acceleration on Leading AI-Optimized FPGAs
by: Taka, Endri, et al.
Published: (2024)
by: Taka, Endri, et al.
Published: (2024)
Interconnect-Aware Logic Resynthesis for Multi-Die FPGAs
by: Wang, Xiaoke, et al.
Published: (2026)
by: Wang, Xiaoke, et al.
Published: (2026)
iDSE: Navigating Design Space Exploration in High-Level Synthesis Using LLMs
by: Li, Runkai, et al.
Published: (2025)
by: Li, Runkai, et al.
Published: (2025)
ForgeHLS: A Large-Scale, Open-Source Dataset for High-Level Synthesis
by: Peng, Zedong, et al.
Published: (2025)
by: Peng, Zedong, et al.
Published: (2025)
H2PIPE: High throughput CNN Inference on FPGAs with High-Bandwidth Memory
by: Doumet, Mario, et al.
Published: (2024)
by: Doumet, Mario, et al.
Published: (2024)
HLS-Eval: A Benchmark and Framework for Evaluating LLMs on High-Level Synthesis Design Tasks
by: Abi-Karam, Stefan, et al.
Published: (2025)
by: Abi-Karam, Stefan, et al.
Published: (2025)
Escaping Flatland: A Placement Flow for Enabling 3D FPGAs
by: Hao, Cong, et al.
Published: (2026)
by: Hao, Cong, et al.
Published: (2026)
Hardware-Efficient Accurate 4-bit Multiplier for Xilinx 7 Series FPGAs
by: Kida, Misaki, et al.
Published: (2025)
by: Kida, Misaki, et al.
Published: (2025)
SA-Kura: An Energy-Efficient Systolic Array Accelerator for Locally-Coupled Kuramoto Drift in Diffusion Sampling
by: Jin, Jeongmin, et al.
Published: (2026)
by: Jin, Jeongmin, et al.
Published: (2026)
RidgeWalker: Perfectly Pipelined Graph Random Walks on FPGAs
by: Tan, Hongshi, et al.
Published: (2026)
by: Tan, Hongshi, et al.
Published: (2026)
HyperSense: Hyperdimensional Intelligent Sensing for Energy-Efficient Sparse Data Processing
by: Yun, Sanggeon, et al.
Published: (2024)
by: Yun, Sanggeon, et al.
Published: (2024)
TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs
by: Qiao, Ye, et al.
Published: (2025)
by: Qiao, Ye, et al.
Published: (2025)
When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference
by: Li, Pu, et al.
Published: (2026)
by: Li, Pu, et al.
Published: (2026)
NLS: Natural-Level Synthesis for Hardware Implementation Through GenAI
by: Yang, Kaiyuan, et al.
Published: (2025)
by: Yang, Kaiyuan, et al.
Published: (2025)
Wavelet Based Frequency Detection Using FPGAs
by: Hill, Caleb, et al.
Published: (2024)
by: Hill, Caleb, et al.
Published: (2024)
Learning to Compare Hardware Designs for High-Level Synthesis
by: Bai, Yunsheng, et al.
Published: (2024)
by: Bai, Yunsheng, et al.
Published: (2024)
Leveraging Compute-in-Memory for Efficient Generative Model Inference in TPUs
by: Zhu, Zhantong, et al.
Published: (2025)
by: Zhu, Zhantong, et al.
Published: (2025)
HLSPilot: LLM-based High-Level Synthesis
by: Xiong, Chenwei, et al.
Published: (2024)
by: Xiong, Chenwei, et al.
Published: (2024)
Dynamic Loop Fusion in High-Level Synthesis
by: Szafarczyk, Robert, et al.
Published: (2025)
by: Szafarczyk, Robert, et al.
Published: (2025)
Theoretical Analysis of the Efficient-Memory Matrix Storage Method for Quantum Emulation Accelerators with Gate Fusion on FPGAs
by: Le, Tran Xuan Hieu, et al.
Published: (2024)
by: Le, Tran Xuan Hieu, et al.
Published: (2024)
PD-Swap: Prefill-Decode Logic Swapping for End-to-End LLM Inference on Edge FPGAs via Dynamic Partial Reconfiguration
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Hardware-Software Co-Design for Event-Driven SNN Deployment on Low-Cost Neuromorphic FPGAs
by: Lee, Jiwoon, et al.
Published: (2026)
by: Lee, Jiwoon, et al.
Published: (2026)
ChatHLS: Towards Systematic Design Automation and Optimization for High-Level Synthesis
by: Li, Runkai, et al.
Published: (2025)
by: Li, Runkai, et al.
Published: (2025)
Similar Items
-
LlamaF: An Efficient Llama2 Architecture Accelerator on Embedded FPGAs
by: Xu, Han, et al.
Published: (2024) -
LUT-LLM: Efficient Large Language Model Inference with Memory-based Computations on FPGAs
by: He, Zifan, et al.
Published: (2025) -
FlightLLM: Efficient Large Language Model Inference with a Complete Mapping Flow on FPGAs
by: Zeng, Shulin, et al.
Published: (2024) -
SWAT: Scalable and Efficient Window Attention-based Transformers Acceleration on FPGAs
by: Bai, Zhenyu, et al.
Published: (2024) -
Energy Efficient LSTM Accelerators for Embedded FPGAs through Parameterised Architecture Design
by: Qian, Chao, et al.
Published: (2026)