PADE: A Predictor-Free Sparse Attention Accelerator via Unified Execution and Stage Fusion
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Huizheng, Wang, Hongbin, Wang, Zichuan, Yue, Zhiheng, Wang, Yang, Li, Chao, Hu, Yang, Yin, Shouyi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Designing Spatial Architectures for Sparse Attention: STAR Accelerator via Cross-Stage Tiling
by: Wang, Huizheng, et al.
Published: (2025)
by: Wang, Huizheng, et al.
Published: (2025)
MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness
by: Wang, Huizheng, et al.
Published: (2025)
by: Wang, Huizheng, et al.
Published: (2025)
SOFA: A Compute-Memory Optimized Sparsity Accelerator via Cross-Stage Coordinated Tiling
by: Wang, Huizheng, et al.
Published: (2024)
by: Wang, Huizheng, et al.
Published: (2024)
TEMP: A Memory Efficient Physical-aware Tensor Partition-Mapping Framework on Wafer-scale Chips
by: Wang, Huizheng, et al.
Published: (2025)
by: Wang, Huizheng, et al.
Published: (2025)
BitStopper: An Efficient Transformer Attention Accelerator via Stage-fusion and Early Termination
by: Wang, Huizheng, et al.
Published: (2025)
by: Wang, Huizheng, et al.
Published: (2025)
From Quarter to All: Accelerating Speculative LLM Decoding via Floating-Point Exponent Remapping and Parameter Sharing
by: Zhao, Yushu, et al.
Published: (2025)
by: Zhao, Yushu, et al.
Published: (2025)
SparseDPD: A Sparse Neural Network-based Digital Predistortion FPGA Accelerator for RF Power Amplifier Linearization
by: Versluis, Manno, et al.
Published: (2025)
by: Versluis, Manno, et al.
Published: (2025)
A High-Throughput Hardware Accelerator for Lempel-Ziv 4 Compression Algorithm
by: Chen, Tao, et al.
Published: (2024)
by: Chen, Tao, et al.
Published: (2024)
Efficient Orchestrated AI Workflows Execution on Scale-out Spatial Architecture
by: Deng, Jinyi, et al.
Published: (2024)
by: Deng, Jinyi, et al.
Published: (2024)
A Digital SRAM-Based Compute-In-Memory Macro for Weight-Stationary Dynamic Matrix Multiplication in Transformer Attention Score Computation
by: Yu, Jianyi, et al.
Published: (2025)
by: Yu, Jianyi, et al.
Published: (2025)
RRAM-Based Bio-Inspired Circuits for Mobile Epileptic Correlation Extraction and Seizure Prediction
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
Pipeline Stage Resolved Timing Characterization of FPGA and ASIC Implementations of a RISC V Processor
by: Darvishi, Mostafa
Published: (2025)
by: Darvishi, Mostafa
Published: (2025)
Channel-Coherence-Adaptive Two-Stage Fully Digital Combining for mmWave MIMO Systems
by: Khorsandmanesh, Yasaman, et al.
Published: (2025)
by: Khorsandmanesh, Yasaman, et al.
Published: (2025)
Lightator: An Optical Near-Sensor Accelerator with Compressive Acquisition Enabling Versatile Image Processing
by: Morsali, Mehrdad, et al.
Published: (2024)
by: Morsali, Mehrdad, et al.
Published: (2024)
Single-Event Upset Analysis of a Systolic Array based Deep Neural Network Accelerator
by: Jonckers, Naïn, et al.
Published: (2024)
by: Jonckers, Naïn, et al.
Published: (2024)
Slimmed optical neural networks with multiplexed neuron sets and a corresponding backpropagation training algorithm
by: Liu, Yi-Feng, et al.
Published: (2023)
by: Liu, Yi-Feng, et al.
Published: (2023)
SnipSnap: A Joint Compression Format and Dataflow Co-Optimization Framework for Efficient Sparse LLM Accelerator Design
by: Wu, Junyi, et al.
Published: (2025)
by: Wu, Junyi, et al.
Published: (2025)
VIKIN: A Reconfigurable Accelerator for KANs and MLPs with Two-Stage Sparsity Support
by: Ou, Wenhui, et al.
Published: (2026)
by: Ou, Wenhui, et al.
Published: (2026)
Aging Aware Adaptive Voltage Scaling for Reliable and Efficient AI Accelerators
by: Xie, Tong, et al.
Published: (2026)
by: Xie, Tong, et al.
Published: (2026)
DEFA: Efficient Deformable Attention Acceleration via Pruning-Assisted Grid-Sampling and Multi-Scale Parallel Processing
by: Xu, Yansong, et al.
Published: (2024)
by: Xu, Yansong, et al.
Published: (2024)
An Efficient Sparse Hardware Accelerator for Spike-Driven Transformer
by: Li, Zhengke, et al.
Published: (2025)
by: Li, Zhengke, et al.
Published: (2025)
Memory-Guided Unified Hardware Accelerator for Mixed-Precision Scientific Computing
by: Wang, Chuanzhen, et al.
Published: (2026)
by: Wang, Chuanzhen, et al.
Published: (2026)
FAST-Prefill: FPGA Accelerated Sparse Attention for Long Context LLM Prefill
by: Jayanth, Rakshith, et al.
Published: (2026)
by: Jayanth, Rakshith, et al.
Published: (2026)
Practical Timing Closure in FPGA and ASIC Designs: Methods, Challenges, and Case Studies
by: Darvishi, Mostafa
Published: (2025)
by: Darvishi, Mostafa
Published: (2025)
Low Complexity Deep Learning Augmented Wireless Channel Estimation for Pilot-Based OFDM on Zynq System on Chip
by: Sharma, Animesh, et al.
Published: (2024)
by: Sharma, Animesh, et al.
Published: (2024)
In-Memory Computing Enabled Deep MIMO Detection to Support Ultra-Low-Latency Communications
by: Ding, Tingyu, et al.
Published: (2025)
by: Ding, Tingyu, et al.
Published: (2025)
Ellora: Exploring Low-Power OFDM-based Radar Processors using Approximate Computing
by: Bhattacharjya, Rajat, et al.
Published: (2023)
by: Bhattacharjya, Rajat, et al.
Published: (2023)
A 0.32 mm$^2$ 100 Mb/s 223 mW ASIC in 22FDX for Joint Jammer Mitigation, Channel Estimation, and SIMO Data Detection
by: Elmiger, Jonas, et al.
Published: (2025)
by: Elmiger, Jonas, et al.
Published: (2025)
FieldHAR: A Fully Integrated End-to-end RTL Framework for Human Activity Recognition with Neural Networks from Heterogeneous Sensors
by: Liu, Mengxi, et al.
Published: (2023)
by: Liu, Mengxi, et al.
Published: (2023)
Spectral Impact of Mismatches in Interleaved ADCs
by: Guichemerre, Jérémy, et al.
Published: (2026)
by: Guichemerre, Jérémy, et al.
Published: (2026)
12-bit Delta-Sigma ADC operating at a temperature of up to 250C in Standard 0.18 $μ$m SOI CMOS
by: Sbrana, Christian, et al.
Published: (2024)
by: Sbrana, Christian, et al.
Published: (2024)
InfiniteEn: A Multi-Source Energy Harvesting System with Load Monitoring Module for Batteryless Internet of Things
by: Puluckul, Priyesh Pappinisseri, et al.
Published: (2024)
by: Puluckul, Priyesh Pappinisseri, et al.
Published: (2024)
Compute SNR-Optimal Analog-to-Digital Converters for Analog In-Memory Computing
by: Kavishwar, Mihir, et al.
Published: (2025)
by: Kavishwar, Mihir, et al.
Published: (2025)
Fixed-Throughput GRAND with FIFO Scheduling
by: Christen, Filippo, et al.
Published: (2025)
by: Christen, Filippo, et al.
Published: (2025)
Decade-Bandwidth RF-Input Pseudo-Doherty Load Modulated Balanced Amplifier using Signal-Flow-Based Phase Alignment Design
by: Gong, Pingzhu, et al.
Published: (2024)
by: Gong, Pingzhu, et al.
Published: (2024)
Design and In-training Optimization of Binary Search ADC for Flexible Classifiers
by: Duarte, Paula Carolina Lozano, et al.
Published: (2024)
by: Duarte, Paula Carolina Lozano, et al.
Published: (2024)
A Custom IC Layout Generation Engine Based on Dynamic Templates and Grids
by: Shin, Taeho, et al.
Published: (2022)
by: Shin, Taeho, et al.
Published: (2022)
Hardware Implementation of Soft Mapper/Demappers in Iterative EP-based Receivers
by: Schilling, Ian Fischer, et al.
Published: (2024)
by: Schilling, Ian Fischer, et al.
Published: (2024)
Automated SAR ADC Sizing Using Analytical Equations
by: Li, Zhongyi, et al.
Published: (2025)
by: Li, Zhongyi, et al.
Published: (2025)
Demonstrator Testbed for Effective Precoding in MEO Multibeam Satellites
by: González-Rios, Jorge L., et al.
Published: (2025)
by: González-Rios, Jorge L., et al.
Published: (2025)
Similar Items
-
Designing Spatial Architectures for Sparse Attention: STAR Accelerator via Cross-Stage Tiling
by: Wang, Huizheng, et al.
Published: (2025) -
MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness
by: Wang, Huizheng, et al.
Published: (2025) -
SOFA: A Compute-Memory Optimized Sparsity Accelerator via Cross-Stage Coordinated Tiling
by: Wang, Huizheng, et al.
Published: (2024) -
TEMP: A Memory Efficient Physical-aware Tensor Partition-Mapping Framework on Wafer-scale Chips
by: Wang, Huizheng, et al.
Published: (2025) -
BitStopper: An Efficient Transformer Attention Accelerator via Stage-fusion and Early Termination
by: Wang, Huizheng, et al.
Published: (2025)