Accelerating Multi-Scale Deformable Attention Using Near-Memory-Processing Architecture
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Huize, Wang, Qinggang, Gao, Bing, Chen, Dan, Huang, Yu, Xin, Xin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hybrid Photonic-digital Accelerator for Attention Mechanism
von: Li, Huize, et al.
Veröffentlicht: (2025)
von: Li, Huize, et al.
Veröffentlicht: (2025)
SCREME: A Scalable Framework for Resilient Memory Design
von: Li, Fan, et al.
Veröffentlicht: (2025)
von: Li, Fan, et al.
Veröffentlicht: (2025)
DEFA: Efficient Deformable Attention Acceleration via Pruning-Assisted Grid-Sampling and Multi-Scale Parallel Processing
von: Xu, Yansong, et al.
Veröffentlicht: (2024)
von: Xu, Yansong, et al.
Veröffentlicht: (2024)
Ironman: Accelerating Oblivious Transfer Extension for Privacy-Preserving AI with Near-Memory Processing
von: Lin, Chenqi, et al.
Veröffentlicht: (2025)
von: Lin, Chenqi, et al.
Veröffentlicht: (2025)
SWAT: Scalable and Efficient Window Attention-based Transformers Acceleration on FPGAs
von: Bai, Zhenyu, et al.
Veröffentlicht: (2024)
von: Bai, Zhenyu, et al.
Veröffentlicht: (2024)
APACHE: A Processing-Near-Memory Architecture for Multi-Scheme Fully Homomorphic Encryption
von: Ding, Lin, et al.
Veröffentlicht: (2024)
von: Ding, Lin, et al.
Veröffentlicht: (2024)
CMD: A Cache-assisted GPU Memory Deduplication Architecture
von: Zhao, Wei, et al.
Veröffentlicht: (2024)
von: Zhao, Wei, et al.
Veröffentlicht: (2024)
PIMCOMP: An End-to-End DNN Compiler for Processing-In-Memory Accelerators
von: Sun, Xiaotian, et al.
Veröffentlicht: (2024)
von: Sun, Xiaotian, et al.
Veröffentlicht: (2024)
End-to-End Transformer Acceleration Through Processing-in-Memory Architectures
von: Yang, Xiaoxuan, et al.
Veröffentlicht: (2025)
von: Yang, Xiaoxuan, et al.
Veröffentlicht: (2025)
FusionCIM: Accelerating LLM Inference with Fusion-Driven Computing-in-Memory Architecture
von: Xuan, Zihao, et al.
Veröffentlicht: (2026)
von: Xuan, Zihao, et al.
Veröffentlicht: (2026)
PUMA: Efficient and Low-Cost Memory Allocation and Alignment Support for Processing-Using-Memory Architectures
von: Oliveira, Geraldo F., et al.
Veröffentlicht: (2024)
von: Oliveira, Geraldo F., et al.
Veröffentlicht: (2024)
PIMSIM-NN: An ISA-based Simulation Framework for Processing-in-Memory Accelerators
von: Wang, Xinyu, et al.
Veröffentlicht: (2024)
von: Wang, Xinyu, et al.
Veröffentlicht: (2024)
AutoRAC: Automated Processing-in-Memory Accelerator Design for Recommender Systems
von: Cheng, Feng, et al.
Veröffentlicht: (2025)
von: Cheng, Feng, et al.
Veröffentlicht: (2025)
NDSEARCH: Accelerating Graph-Traversal-Based Approximate Nearest Neighbor Search through Near Data Processing
von: Wang, Yitu, et al.
Veröffentlicht: (2023)
von: Wang, Yitu, et al.
Veröffentlicht: (2023)
Be CIM or Be Memory: A Dual-mode-aware DNN Compiler for CIM Accelerators
von: Zhao, Shixin, et al.
Veröffentlicht: (2025)
von: Zhao, Shixin, et al.
Veröffentlicht: (2025)
UpANNS: Enhancing Billion-Scale ANNS Efficiency with Real-World PIM Architecture
von: Chen, Sitian, et al.
Veröffentlicht: (2024)
von: Chen, Sitian, et al.
Veröffentlicht: (2024)
Hybrid SLC-MLC RRAM Mixed-Signal Processing-in-Memory Architecture for Transformer Acceleration via Gradient Redistribution
von: Song, Chang Eun, et al.
Veröffentlicht: (2025)
von: Song, Chang Eun, et al.
Veröffentlicht: (2025)
Generalized Ping-Pong: Off-Chip Memory Bandwidth Centric Pipelining Strategy for Processing-In-Memory Accelerators
von: Wang, Ruibao, et al.
Veröffentlicht: (2024)
von: Wang, Ruibao, et al.
Veröffentlicht: (2024)
PIM-GPT: A Hybrid Process-in-Memory Accelerator for Autoregressive Transformers
von: Wu, Yuting, et al.
Veröffentlicht: (2023)
von: Wu, Yuting, et al.
Veröffentlicht: (2023)
AccelCIM: Systematic Dataflow Exploration for SRAM Compute-in-Memory Accelerator
von: Xue, Chenhao, et al.
Veröffentlicht: (2026)
von: Xue, Chenhao, et al.
Veröffentlicht: (2026)
Scalable and RISC-V Programmable Near-Memory Computing Architectures for Edge Nodes
von: Caon, Michele, et al.
Veröffentlicht: (2024)
von: Caon, Michele, et al.
Veröffentlicht: (2024)
HPIM: Heterogeneous Processing-In-Memory-based Accelerator for Large Language Models Inference
von: Duan, Cenlin, et al.
Veröffentlicht: (2025)
von: Duan, Cenlin, et al.
Veröffentlicht: (2025)
A Switch-Centric In-Network Architecture for Accelerating LLM Inference in Shared-Memory Network
von: Jiang, Aojie, et al.
Veröffentlicht: (2026)
von: Jiang, Aojie, et al.
Veröffentlicht: (2026)
FlexMem: High-Parallel Near-Memory Architecture for Flexible Dataflow in Fully Homomorphic Encryption
von: Shi, Shangyi, et al.
Veröffentlicht: (2025)
von: Shi, Shangyi, et al.
Veröffentlicht: (2025)
PRIMAL: Processing-In-Memory Based Low-Rank Adaptation for LLM Inference Accelerator
von: Chong, Yue Jiet, et al.
Veröffentlicht: (2026)
von: Chong, Yue Jiet, et al.
Veröffentlicht: (2026)
Near-Memory Architecture for Threshold-Ordinal Surface-Based Corner Detection of Event Cameras
von: Shang, Hongyang, et al.
Veröffentlicht: (2025)
von: Shang, Hongyang, et al.
Veröffentlicht: (2025)
Co-Designing Graph-based Approximate Nearest Neighbor Search at Billion Scale for Processing-in-Memory
von: Chen, Sitian, et al.
Veröffentlicht: (2026)
von: Chen, Sitian, et al.
Veröffentlicht: (2026)
PIM-malloc: A Fast and Scalable Dynamic Memory Allocator for Processing-In-Memory (PIM) Architectures
von: Lee, Dongjae, et al.
Veröffentlicht: (2025)
von: Lee, Dongjae, et al.
Veröffentlicht: (2025)
EPIM: Efficient Processing-In-Memory Accelerators based on Epitome
von: Wang, Chenyu, et al.
Veröffentlicht: (2023)
von: Wang, Chenyu, et al.
Veröffentlicht: (2023)
PIMSYN: Synthesizing Processing-in-memory CNN Accelerators
von: Li, Wanqian, et al.
Veröffentlicht: (2024)
von: Li, Wanqian, et al.
Veröffentlicht: (2024)
Low-overhead General-purpose Near-Data Processing in CXL Memory Expanders
von: Ham, Hyungkyu, et al.
Veröffentlicht: (2024)
von: Ham, Hyungkyu, et al.
Veröffentlicht: (2024)
SigDLA: A Deep Learning Accelerator Extension for Signal Processing
von: Fu, Fangfa, et al.
Veröffentlicht: (2024)
von: Fu, Fangfa, et al.
Veröffentlicht: (2024)
Efficient Sparse Processing-in-Memory Architecture (ESPIM) for Machine Learning Inference
von: He, Mingxuan, et al.
Veröffentlicht: (2024)
von: He, Mingxuan, et al.
Veröffentlicht: (2024)
A System Architecture for Low Latency Multiprogramming Quantum Computing
von: Zhao, Yilun, et al.
Veröffentlicht: (2026)
von: Zhao, Yilun, et al.
Veröffentlicht: (2026)
Accelerating Time Series Analysis via Processing using Non-Volatile Memories
von: Fernandez, Ivan, et al.
Veröffentlicht: (2022)
von: Fernandez, Ivan, et al.
Veröffentlicht: (2022)
In-Pipeline Integration of Digital In-Memory-Computing into RISC-V Vector Architecture to Accelerate Deep Learning
von: Spagnolo, Tommaso, et al.
Veröffentlicht: (2026)
von: Spagnolo, Tommaso, et al.
Veröffentlicht: (2026)
DaPPA: A Data-Parallel Programming Framework for Processing-in-Memory Architectures
von: Oliveira, Geraldo F., et al.
Veröffentlicht: (2023)
von: Oliveira, Geraldo F., et al.
Veröffentlicht: (2023)
Accelerating Seed Location Filtering in DNA Read Mapping Using a Commercial Compute-in-SRAM Architecture
von: Golden, Courtney, et al.
Veröffentlicht: (2024)
von: Golden, Courtney, et al.
Veröffentlicht: (2024)
SwiftKV: An Edge-Oriented Attention Algorithm and Multi-Head Accelerator for Fast, Efficient LLM Decoding
von: Zhang, Junming, et al.
Veröffentlicht: (2026)
von: Zhang, Junming, et al.
Veröffentlicht: (2026)
Hemlet: A Heterogeneous Compute-in-Memory Chiplet Architecture for Vision Transformers with Group-Level Parallelism
von: Wang, Cong, et al.
Veröffentlicht: (2025)
von: Wang, Cong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Hybrid Photonic-digital Accelerator for Attention Mechanism
von: Li, Huize, et al.
Veröffentlicht: (2025) -
SCREME: A Scalable Framework for Resilient Memory Design
von: Li, Fan, et al.
Veröffentlicht: (2025) -
DEFA: Efficient Deformable Attention Acceleration via Pruning-Assisted Grid-Sampling and Multi-Scale Parallel Processing
von: Xu, Yansong, et al.
Veröffentlicht: (2024) -
Ironman: Accelerating Oblivious Transfer Extension for Privacy-Preserving AI with Near-Memory Processing
von: Lin, Chenqi, et al.
Veröffentlicht: (2025) -
SWAT: Scalable and Efficient Window Attention-based Transformers Acceleration on FPGAs
von: Bai, Zhenyu, et al.
Veröffentlicht: (2024)