DEFA: Efficient Deformable Attention Acceleration via Pruning-Assisted Grid-Sampling and Multi-Scale Parallel Processing
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Yansong, Lyu, Dongxu, Li, Zhenyu, Wang, Zilong, Chen, Yuzhou, Wang, Gang, Wang, Zhican, Li, Haomin, He, Guanghui |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SD-Acc: Accelerating Stable Diffusion through Phase-aware Sampling and Hardware Co-Optimizations
by: Wang, Zhican, et al.
Published: (2025)
by: Wang, Zhican, et al.
Published: (2025)
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator
by: Wang, Zhican, et al.
Published: (2025)
by: Wang, Zhican, et al.
Published: (2025)
Accelerating Multi-Scale Deformable Attention Using Near-Memory-Processing Architecture
by: Li, Huize, et al.
Published: (2026)
by: Li, Huize, et al.
Published: (2026)
PIMSYN: Synthesizing Processing-in-memory CNN Accelerators
by: Li, Wanqian, et al.
Published: (2024)
by: Li, Wanqian, et al.
Published: (2024)
GEMM-GS: Accelerating 3D Gaussian Splatting on Tensor Cores with GEMM-Compatible Blending
by: Li, Haomin, et al.
Published: (2026)
by: Li, Haomin, et al.
Published: (2026)
Aging Aware Adaptive Voltage Scaling for Reliable and Efficient AI Accelerators
by: Xie, Tong, et al.
Published: (2026)
by: Xie, Tong, et al.
Published: (2026)
PIMCOMP: An End-to-End DNN Compiler for Processing-In-Memory Accelerators
by: Sun, Xiaotian, et al.
Published: (2024)
by: Sun, Xiaotian, et al.
Published: (2024)
HiHGNN: Accelerating HGNNs through Parallelism and Data Reusability Exploitation
by: Xue, Runzhen, et al.
Published: (2023)
by: Xue, Runzhen, et al.
Published: (2023)
LaMoS: Enabling Efficient Large Number Modular Multiplication through SRAM-based CiM Acceleration
by: Li, Haomin, et al.
Published: (2025)
by: Li, Haomin, et al.
Published: (2025)
SWAT: Scalable and Efficient Window Attention-based Transformers Acceleration on FPGAs
by: Bai, Zhenyu, et al.
Published: (2024)
by: Bai, Zhenyu, et al.
Published: (2024)
Hybrid Photonic-digital Accelerator for Attention Mechanism
by: Li, Huize, et al.
Published: (2025)
by: Li, Huize, et al.
Published: (2025)
SigDLA: A Deep Learning Accelerator Extension for Signal Processing
by: Fu, Fangfa, et al.
Published: (2024)
by: Fu, Fangfa, et al.
Published: (2024)
Ironman: Accelerating Oblivious Transfer Extension for Privacy-Preserving AI with Near-Memory Processing
by: Lin, Chenqi, et al.
Published: (2025)
by: Lin, Chenqi, et al.
Published: (2025)
NDSEARCH: Accelerating Graph-Traversal-Based Approximate Nearest Neighbor Search through Near Data Processing
by: Wang, Yitu, et al.
Published: (2023)
by: Wang, Yitu, et al.
Published: (2023)
Reconfigurable Digital RRAM Logic Enables In-Situ Pruning and Learning for Edge AI
by: Wang, Songqi, et al.
Published: (2025)
by: Wang, Songqi, et al.
Published: (2025)
AutoRAC: Automated Processing-in-Memory Accelerator Design for Recommender Systems
by: Cheng, Feng, et al.
Published: (2025)
by: Cheng, Feng, et al.
Published: (2025)
PIM-GPT: A Hybrid Process-in-Memory Accelerator for Autoregressive Transformers
by: Wu, Yuting, et al.
Published: (2023)
by: Wu, Yuting, et al.
Published: (2023)
HPIM: Heterogeneous Processing-In-Memory-based Accelerator for Large Language Models Inference
by: Duan, Cenlin, et al.
Published: (2025)
by: Duan, Cenlin, et al.
Published: (2025)
PIMSIM-NN: An ISA-based Simulation Framework for Processing-in-Memory Accelerators
by: Wang, Xinyu, et al.
Published: (2024)
by: Wang, Xinyu, et al.
Published: (2024)
ISAAC: Intelligent, Scalable, Agile, and Accelerated CPU Verification via LLM-aided FPGA Parallelism
by: Sun, Jialin, et al.
Published: (2025)
by: Sun, Jialin, et al.
Published: (2025)
Generalized Ping-Pong: Off-Chip Memory Bandwidth Centric Pipelining Strategy for Processing-In-Memory Accelerators
by: Wang, Ruibao, et al.
Published: (2024)
by: Wang, Ruibao, et al.
Published: (2024)
Titanus: Enabling KV Cache Pruning and Quantization On-the-Fly for LLM Acceleration
by: Chen, Peilin, et al.
Published: (2025)
by: Chen, Peilin, et al.
Published: (2025)
PRIMAL: Processing-In-Memory Based Low-Rank Adaptation for LLM Inference Accelerator
by: Chong, Yue Jiet, et al.
Published: (2026)
by: Chong, Yue Jiet, et al.
Published: (2026)
PADE: A Predictor-Free Sparse Attention Accelerator via Unified Execution and Stage Fusion
by: Wang, Huizheng, et al.
Published: (2025)
by: Wang, Huizheng, et al.
Published: (2025)
Be CIM or Be Memory: A Dual-mode-aware DNN Compiler for CIM Accelerators
by: Zhao, Shixin, et al.
Published: (2025)
by: Zhao, Shixin, et al.
Published: (2025)
ASDR: Exploiting Adaptive Sampling and Data Reuse for CIM-based Instant Neural Rendering
by: Liu, Fangxin, et al.
Published: (2025)
by: Liu, Fangxin, et al.
Published: (2025)
FASE: FPGA-Assisted Syscall Emulation for Rapid End-to-End Processor Performance Validation
by: Meng, Chengzhen, et al.
Published: (2025)
by: Meng, Chengzhen, et al.
Published: (2025)
Multilayer Dataflow: Orchestrate Butterfly Sparsity to Accelerate Attention Computation
by: Wu, Haibin, et al.
Published: (2024)
by: Wu, Haibin, et al.
Published: (2024)
Fast-OverlaPIM: A Fast Overlap-driven Mapping Framework for Processing In-Memory Neural Network Acceleration
by: Wang, Xuan, et al.
Published: (2024)
by: Wang, Xuan, et al.
Published: (2024)
HiMA: Hierarchical Quantum Microarchitecture for Qubit-Scaling and Quantum Process-Level Parallelism
by: Zhou, Qi, et al.
Published: (2024)
by: Zhou, Qi, et al.
Published: (2024)
STAR: An Efficient Softmax Engine for Attention Model with RRAM Crossbar
by: Zhai, Yifeng, et al.
Published: (2024)
by: Zhai, Yifeng, et al.
Published: (2024)
The Quest for Reliable AI Accelerators: Cross-Layer Evaluation and Design Optimization
by: Li, Meng, et al.
Published: (2026)
by: Li, Meng, et al.
Published: (2026)
GTA: a new General Tensor Accelerator with Better Area Efficiency and Data Reuse
by: Ai, Chenyang, et al.
Published: (2024)
by: Ai, Chenyang, et al.
Published: (2024)
Manticore: Hardware-Accelerated RTL Simulation with Static Bulk-Synchronous Parallelism
by: Emami, Mahyar, et al.
Published: (2023)
by: Emami, Mahyar, et al.
Published: (2023)
Graphitron: A Domain Specific Language for FPGA-based Graph Processing Accelerator Generation
by: Zhang, Xinmiao, et al.
Published: (2024)
by: Zhang, Xinmiao, et al.
Published: (2024)
An Efficient Sparse Hardware Accelerator for Spike-Driven Transformer
by: Li, Zhengke, et al.
Published: (2025)
by: Li, Zhengke, et al.
Published: (2025)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
FuseFPS: Accelerating Farthest Point Sampling with Fusing KD-tree Construction for Point Clouds
by: Han, Meng, et al.
Published: (2023)
by: Han, Meng, et al.
Published: (2023)
GPU-Accelerated Optimization Solver for Unit Commitment in Large-Scale Power Grids
by: Sharadga, Hussein, et al.
Published: (2025)
by: Sharadga, Hussein, et al.
Published: (2025)
GOMA: Geometrically Optimal Mapping via Analytical Modeling for Spatial Accelerators
by: Yang, Wulve, et al.
Published: (2026)
by: Yang, Wulve, et al.
Published: (2026)
Similar Items
-
SD-Acc: Accelerating Stable Diffusion through Phase-aware Sampling and Hardware Co-Optimizations
by: Wang, Zhican, et al.
Published: (2025) -
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator
by: Wang, Zhican, et al.
Published: (2025) -
Accelerating Multi-Scale Deformable Attention Using Near-Memory-Processing Architecture
by: Li, Huize, et al.
Published: (2026) -
PIMSYN: Synthesizing Processing-in-memory CNN Accelerators
by: Li, Wanqian, et al.
Published: (2024) -
GEMM-GS: Accelerating 3D Gaussian Splatting on Tensor Cores with GEMM-Compatible Blending
by: Li, Haomin, et al.
Published: (2026)