A Persistent-State Dataflow Accelerator for Memory-Bound Linear Attention Decode on FPGA
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gupta, Neelesh, Wang, Peter, Kannan, Rajgopal, Prasanna, Viktor K. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PaCKD: Pattern-Clustered Knowledge Distillation for Compressing Memory Access Prediction Models
von: Gupta, Neelesh, et al.
Veröffentlicht: (2024)
von: Gupta, Neelesh, et al.
Veröffentlicht: (2024)
Attention, Distillation, and Tabularization: Towards Practical Neural Network-Based Prefetching
von: Zhang, Pengmiao, et al.
Veröffentlicht: (2023)
von: Zhang, Pengmiao, et al.
Veröffentlicht: (2023)
Enabling Long FFT Convolutions on Memory-Constrained FPGAs via Chunking
von: Wang, Peter, et al.
Veröffentlicht: (2025)
von: Wang, Peter, et al.
Veröffentlicht: (2025)
FAME: FPGA Acceleration of Secure Matrix Multiplication with Homomorphic Encryption
von: Xu, Zhihan, et al.
Veröffentlicht: (2025)
von: Xu, Zhihan, et al.
Veröffentlicht: (2025)
Sparse MTTKRP Acceleration for Tensor Decomposition on GPU
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2024)
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2024)
FAST-Prefill: FPGA Accelerated Sparse Attention for Long Context LLM Prefill
von: Jayanth, Rakshith, et al.
Veröffentlicht: (2026)
von: Jayanth, Rakshith, et al.
Veröffentlicht: (2026)
Accelerating ViT Inference on FPGA through Static and Dynamic Pruning
von: Parikh, Dhruv, et al.
Veröffentlicht: (2024)
von: Parikh, Dhruv, et al.
Veröffentlicht: (2024)
Accelerating Recommender Model ETL with a Streaming FPGA-GPU Dataflow
von: Zhu, Yu, et al.
Veröffentlicht: (2025)
von: Zhu, Yu, et al.
Veröffentlicht: (2025)
HASS: Hardware-Aware Sparsity Search for Dataflow DNN Accelerator
von: Yu, Zhewen, et al.
Veröffentlicht: (2024)
von: Yu, Zhewen, et al.
Veröffentlicht: (2024)
DG-RePlAce: A Dataflow-Driven GPU-Accelerated Analytical Global Placement Framework for Machine Learning Accelerators
von: Kahng, Andrew B., et al.
Veröffentlicht: (2024)
von: Kahng, Andrew B., et al.
Veröffentlicht: (2024)
Low Power Vision Transformer Accelerator with Hardware-Aware Pruning and Optimized Dataflow
von: Hsiung, Ching-Lin, et al.
Veröffentlicht: (2025)
von: Hsiung, Ching-Lin, et al.
Veröffentlicht: (2025)
VTR: An Optimized Vision Transformer for SAR ATR Acceleration on FPGA
von: Wickramasinghe, Sachini, et al.
Veröffentlicht: (2024)
von: Wickramasinghe, Sachini, et al.
Veröffentlicht: (2024)
An FPGA-Based Reconfigurable Accelerator for Convolution-Transformer Hybrid EfficientViT
von: Shao, Haikuo, et al.
Veröffentlicht: (2024)
von: Shao, Haikuo, et al.
Veröffentlicht: (2024)
An FPGA-Based Accelerator Enabling Efficient Support for CNNs with Arbitrary Kernel Sizes
von: Wang, Miaoxin, et al.
Veröffentlicht: (2024)
von: Wang, Miaoxin, et al.
Veröffentlicht: (2024)
Memory-Efficient FPGA Implementation of Stochastic Simulated Annealing
von: Shin, Duckgyu, et al.
Veröffentlicht: (2026)
von: Shin, Duckgyu, et al.
Veröffentlicht: (2026)
A High-Throughput FPGA Accelerator for Lightweight CNNs With Balanced Dataflow
von: Zhao, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Zhao, Zhiyuan, et al.
Veröffentlicht: (2024)
Dato: A Task-Based Programming Model for Dataflow Accelerators
von: Fang, Shihan, et al.
Veröffentlicht: (2025)
von: Fang, Shihan, et al.
Veröffentlicht: (2025)
Efficient and Accurate Graph Classification with Hyperdimensional Computing on FPGA
von: Arockiaraj, Jebacyril, et al.
Veröffentlicht: (2025)
von: Arockiaraj, Jebacyril, et al.
Veröffentlicht: (2025)
Token-Picker: Accelerating Attention in Text Generation with Minimized Memory Transfer via Probability Estimation
von: Park, Junyoung, et al.
Veröffentlicht: (2024)
von: Park, Junyoung, et al.
Veröffentlicht: (2024)
NSFlow: An End-to-End FPGA Framework with Scalable Dataflow Architecture for Neuro-Symbolic AI
von: Yang, Hanchen, et al.
Veröffentlicht: (2025)
von: Yang, Hanchen, et al.
Veröffentlicht: (2025)
M100: An Orchestrated Dataflow Architecture Powering General AI Computing
von: Xie, Yan, et al.
Veröffentlicht: (2026)
von: Xie, Yan, et al.
Veröffentlicht: (2026)
SIRA: Scaled-Integer Range Analysis for Optimizing FPGA Dataflow Neural Network Accelerators
von: Umuroglu, Yaman, et al.
Veröffentlicht: (2025)
von: Umuroglu, Yaman, et al.
Veröffentlicht: (2025)
EPIM: Efficient Processing-In-Memory Accelerators based on Epitome
von: Wang, Chenyu, et al.
Veröffentlicht: (2023)
von: Wang, Chenyu, et al.
Veröffentlicht: (2023)
MEADOW: Memory-efficient Dataflow and Data Packing for Low Power Edge LLMs
von: Moitra, Abhishek, et al.
Veröffentlicht: (2025)
von: Moitra, Abhishek, et al.
Veröffentlicht: (2025)
A Data-Driven Approach to Dataflow-Aware Online Scheduling for Graph Neural Network Inference
von: Puigdemont, Pol, et al.
Veröffentlicht: (2024)
von: Puigdemont, Pol, et al.
Veröffentlicht: (2024)
Revealing CNN Architectures via Side-Channel Analysis in Dataflow-based Inference Accelerators
von: Weerasena, Hansika, et al.
Veröffentlicht: (2023)
von: Weerasena, Hansika, et al.
Veröffentlicht: (2023)
TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs
von: Qiao, Ye, et al.
Veröffentlicht: (2025)
von: Qiao, Ye, et al.
Veröffentlicht: (2025)
EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture
von: Duan, Bowen, et al.
Veröffentlicht: (2026)
von: Duan, Bowen, et al.
Veröffentlicht: (2026)
Multilayer Dataflow: Orchestrate Butterfly Sparsity to Accelerate Attention Computation
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
Exploring FPGA designs for MX and beyond
von: Samson, Ebby, et al.
Veröffentlicht: (2024)
von: Samson, Ebby, et al.
Veröffentlicht: (2024)
Efficient In-Memory Acceleration of Sparse Block Diagonal LLMs
von: de Lima, João Paulo Cardoso, et al.
Veröffentlicht: (2025)
von: de Lima, João Paulo Cardoso, et al.
Veröffentlicht: (2025)
AccelCIM: Systematic Dataflow Exploration for SRAM Compute-in-Memory Accelerator
von: Xue, Chenhao, et al.
Veröffentlicht: (2026)
von: Xue, Chenhao, et al.
Veröffentlicht: (2026)
FPGA-based Acceleration for Convolutional Neural Networks: A Comprehensive Review
von: Jiang, Junye, et al.
Veröffentlicht: (2025)
von: Jiang, Junye, et al.
Veröffentlicht: (2025)
Efficient VQ-QAT and Mixed Vector/Linear quantized Neural Networks
von: Gou, Terry, et al.
Veröffentlicht: (2026)
von: Gou, Terry, et al.
Veröffentlicht: (2026)
End-to-End Transformer Acceleration Through Processing-in-Memory Architectures
von: Yang, Xiaoxuan, et al.
Veröffentlicht: (2025)
von: Yang, Xiaoxuan, et al.
Veröffentlicht: (2025)
ITA: An Energy-Efficient Attention and Softmax Accelerator for Quantized Transformers
von: İslamoğlu, Gamze, et al.
Veröffentlicht: (2023)
von: İslamoğlu, Gamze, et al.
Veröffentlicht: (2023)
Memory Is All You Need: An Overview of Compute-in-Memory Architectures for Accelerating Large Language Model Inference
von: Wolters, Christopher, et al.
Veröffentlicht: (2024)
von: Wolters, Christopher, et al.
Veröffentlicht: (2024)
Pimba: A Processing-in-Memory Acceleration for Post-Transformer Large Language Model Serving
von: Kim, Wonung, et al.
Veröffentlicht: (2025)
von: Kim, Wonung, et al.
Veröffentlicht: (2025)
Ultra Memory-Efficient On-FPGA Training of Transformers via Tensor-Compressed Optimization
von: Tian, Jiayi, et al.
Veröffentlicht: (2025)
von: Tian, Jiayi, et al.
Veröffentlicht: (2025)
Exploiting temporal parallelism for LSTM Autoencoder acceleration on FPGA
von: Leftheriotis, Aimilios, et al.
Veröffentlicht: (2026)
von: Leftheriotis, Aimilios, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
PaCKD: Pattern-Clustered Knowledge Distillation for Compressing Memory Access Prediction Models
von: Gupta, Neelesh, et al.
Veröffentlicht: (2024) -
Attention, Distillation, and Tabularization: Towards Practical Neural Network-Based Prefetching
von: Zhang, Pengmiao, et al.
Veröffentlicht: (2023) -
Enabling Long FFT Convolutions on Memory-Constrained FPGAs via Chunking
von: Wang, Peter, et al.
Veröffentlicht: (2025) -
FAME: FPGA Acceleration of Secure Matrix Multiplication with Homomorphic Encryption
von: Xu, Zhihan, et al.
Veröffentlicht: (2025) -
Sparse MTTKRP Acceleration for Tensor Decomposition on GPU
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2024)