SWAT: Scalable and Efficient Window Attention-based Transformers Acceleration on FPGAs
Fuente:
arXiv
Saved in:
| Main Authors: | Bai, Zhenyu, Dangi, Pranav, Li, Huize, Mitra, Tulika |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hybrid Photonic-digital Accelerator for Attention Mechanism
by: Li, Huize, et al.
Published: (2025)
by: Li, Huize, et al.
Published: (2025)
A Data-Driven Dynamic Execution Orchestration Architecture
by: Bai, Zhenyu, et al.
Published: (2026)
by: Bai, Zhenyu, et al.
Published: (2026)
FAMOUS: Flexible Accelerator for the Attention Mechanism of Transformer on UltraScale+ FPGAs
by: Kabir, Ehsan, et al.
Published: (2024)
by: Kabir, Ehsan, et al.
Published: (2024)
TerEffic: Highly Efficient Ternary LLM Inference on FPGA
by: Yin, Chenyang, et al.
Published: (2025)
by: Yin, Chenyang, et al.
Published: (2025)
Enhancing CGRA Efficiency Through Aligned Compute and Communication Provisioning
by: Li, Zhaoying, et al.
Published: (2024)
by: Li, Zhaoying, et al.
Published: (2024)
Sustainable Hardware Specialization
by: Dangi, Pranav, et al.
Published: (2024)
by: Dangi, Pranav, et al.
Published: (2024)
Nexus Machine: An Active Message Inspired Reconfigurable Architecture for Irregular Workloads
by: Juneja, Rohan, et al.
Published: (2025)
by: Juneja, Rohan, et al.
Published: (2025)
HALO: Hardware-aware quantization with low critical-path-delay weights for LLM acceleration
by: Juneja, Rohan, et al.
Published: (2025)
by: Juneja, Rohan, et al.
Published: (2025)
Data-aware Dynamic Execution of Irregular Workloads on Heterogeneous Systems
by: Bai, Zhenyu, et al.
Published: (2025)
by: Bai, Zhenyu, et al.
Published: (2025)
LUT-LLM: Efficient Large Language Model Inference with Memory-based Computations on FPGAs
by: He, Zifan, et al.
Published: (2025)
by: He, Zifan, et al.
Published: (2025)
Building an Open CGRA Ecosystem for Agile Innovation
by: Juneja, Rohan, et al.
Published: (2025)
by: Juneja, Rohan, et al.
Published: (2025)
FlightLLM: Efficient Large Language Model Inference with a Complete Mapping Flow on FPGAs
by: Zeng, Shulin, et al.
Published: (2024)
by: Zeng, Shulin, et al.
Published: (2024)
TReX- Reusing Vision Transformer's Attention for Efficient Xbar-based Computing
by: Moitra, Abhishek, et al.
Published: (2024)
by: Moitra, Abhishek, et al.
Published: (2024)
BETA: Binarized Energy-Efficient Transformer Accelerator at the Edge
by: Ji, Yuhao, et al.
Published: (2024)
by: Ji, Yuhao, et al.
Published: (2024)
Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding
by: Fan, Wang, et al.
Published: (2026)
by: Fan, Wang, et al.
Published: (2026)
Resource Utilization of Differentiable Logic Gate Networks Deployed on FPGAs
by: Wormald, Stephen, et al.
Published: (2026)
by: Wormald, Stephen, et al.
Published: (2026)
Shedding the Bits: Pushing the Boundaries of Quantization with Minifloats on FPGAs
by: Aggarwal, Shivam, et al.
Published: (2023)
by: Aggarwal, Shivam, et al.
Published: (2023)
Hardware/Software Co-Design of RISC-V Extensions for Accelerating Sparse DNNs on FPGAs
by: Sabih, Muhammad, et al.
Published: (2025)
by: Sabih, Muhammad, et al.
Published: (2025)
REASON: Accelerating Probabilistic Logical Reasoning for Scalable Neuro-Symbolic Intelligence
by: Wan, Zishen, et al.
Published: (2026)
by: Wan, Zishen, et al.
Published: (2026)
ProactivePIM: Accelerating Weight-Sharing Embedding Layer with PIM for Scalable Recommendation System
by: Kim, Youngsuk, et al.
Published: (2024)
by: Kim, Youngsuk, et al.
Published: (2024)
Embedded FPGA Acceleration of Brain-Like Neural Networks: Online Learning to Scalable Inference
by: Hafiz, Muhammad Ihsan Al, et al.
Published: (2025)
by: Hafiz, Muhammad Ihsan Al, et al.
Published: (2025)
HLSTransform: Energy-Efficient Llama 2 Inference on FPGAs Via High Level Synthesis
by: He, Andy, et al.
Published: (2024)
by: He, Andy, et al.
Published: (2024)
AttentionLego: An Open-Source Building Block For Spatially-Scalable Large Language Model Accelerator With Processing-In-Memory Technology
by: Cong, Rongqing, et al.
Published: (2024)
by: Cong, Rongqing, et al.
Published: (2024)
Accelerating Multi-Scale Deformable Attention Using Near-Memory-Processing Architecture
by: Li, Huize, et al.
Published: (2026)
by: Li, Huize, et al.
Published: (2026)
A Configurable and Efficient Memory Hierarchy for Neural Network Hardware Accelerator
by: Bause, Oliver, et al.
Published: (2024)
by: Bause, Oliver, et al.
Published: (2024)
TSB: Tiny Shared Block for Efficient DNN Deployment on NVCIM Accelerators
by: Qin, Yifan, et al.
Published: (2024)
by: Qin, Yifan, et al.
Published: (2024)
KAN-SAs: Efficient Acceleration of Kolmogorov-Arnold Networks on Systolic Arrays
by: Errabii, Sohaib, et al.
Published: (2025)
by: Errabii, Sohaib, et al.
Published: (2025)
Carbon-Efficient 3D DNN Acceleration: Optimizing Performance and Sustainability
by: Panteleaki, Aikaterini Maria, et al.
Published: (2025)
by: Panteleaki, Aikaterini Maria, et al.
Published: (2025)
M$^2$-ViT: Accelerating Hybrid Vision Transformers with Two-Level Mixed Quantization
by: Liang, Yanbiao, et al.
Published: (2024)
by: Liang, Yanbiao, et al.
Published: (2024)
LlamaF: An Efficient Llama2 Architecture Accelerator on Embedded FPGAs
by: Xu, Han, et al.
Published: (2024)
by: Xu, Han, et al.
Published: (2024)
Towards Efficient IMC Accelerator Design Through Joint Hardware-Workload Co-optimization
by: Krestinskaya, Olga, et al.
Published: (2024)
by: Krestinskaya, Olga, et al.
Published: (2024)
FastMamba: A High-Speed and Efficient Mamba Accelerator on FPGA with Accurate Quantization
by: Wang, Aotao, et al.
Published: (2025)
by: Wang, Aotao, et al.
Published: (2025)
ApproXAI: Energy-Efficient Hardware Acceleration of Explainable AI using Approximate Computing
by: Siddique, Ayesha, et al.
Published: (2025)
by: Siddique, Ayesha, et al.
Published: (2025)
SystolicAttention: Fusing FlashAttention within a Single Systolic Array
by: Lin, Jiawei, et al.
Published: (2025)
by: Lin, Jiawei, et al.
Published: (2025)
ROMA: a Read-Only-Memory-based Accelerator for QLoRA-based On-Device LLM
by: Wang, Wenqiang, et al.
Published: (2025)
by: Wang, Wenqiang, et al.
Published: (2025)
Runtime Tunable Tsetlin Machines for Edge Inference on eFPGAs
by: Rahman, Tousif, et al.
Published: (2025)
by: Rahman, Tousif, et al.
Published: (2025)
Accelerating MRI Uncertainty Estimation with Mask-based Bayesian Neural Network
by: Zhang, Zehuan, et al.
Published: (2024)
by: Zhang, Zehuan, et al.
Published: (2024)
SALSA: Simulated Annealing based Loop-Ordering Scheduler for DNN Accelerators
by: Jung, Victor J. B., et al.
Published: (2023)
by: Jung, Victor J. B., et al.
Published: (2023)
Accelerating Boolean Constraint Propagation for Efficient SAT-Solving on FPGAs
by: Govindasamy, Hariprasadh, et al.
Published: (2024)
by: Govindasamy, Hariprasadh, et al.
Published: (2024)
Efficient Approaches for GEMM Acceleration on Leading AI-Optimized FPGAs
by: Taka, Endri, et al.
Published: (2024)
by: Taka, Endri, et al.
Published: (2024)
Similar Items
-
Hybrid Photonic-digital Accelerator for Attention Mechanism
by: Li, Huize, et al.
Published: (2025) -
A Data-Driven Dynamic Execution Orchestration Architecture
by: Bai, Zhenyu, et al.
Published: (2026) -
FAMOUS: Flexible Accelerator for the Attention Mechanism of Transformer on UltraScale+ FPGAs
by: Kabir, Ehsan, et al.
Published: (2024) -
TerEffic: Highly Efficient Ternary LLM Inference on FPGA
by: Yin, Chenyang, et al.
Published: (2025) -
Enhancing CGRA Efficiency Through Aligned Compute and Communication Provisioning
by: Li, Zhaoying, et al.
Published: (2024)