Designing Spatial Architectures for Sparse Attention: STAR Accelerator via Cross-Stage Tiling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Huizheng, Wei, Taiquan, Wang, Hongbin, Wang, Zichuan, Tang, Xinru, Yue, Zhiheng, Wei, Shaojun, Hu, Yang, Yin, Shouyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PADE: A Predictor-Free Sparse Attention Accelerator via Unified Execution and Stage Fusion
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
SOFA: A Compute-Memory Optimized Sparsity Accelerator via Cross-Stage Coordinated Tiling
von: Wang, Huizheng, et al.
Veröffentlicht: (2024)
von: Wang, Huizheng, et al.
Veröffentlicht: (2024)
TEMP: A Memory Efficient Physical-aware Tensor Partition-Mapping Framework on Wafer-scale Chips
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
Efficient Orchestrated AI Workflows Execution on Scale-out Spatial Architecture
von: Deng, Jinyi, et al.
Veröffentlicht: (2024)
von: Deng, Jinyi, et al.
Veröffentlicht: (2024)
BitStopper: An Efficient Transformer Attention Accelerator via Stage-fusion and Early Termination
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
From Quarter to All: Accelerating Speculative LLM Decoding via Floating-Point Exponent Remapping and Parameter Sharing
von: Zhao, Yushu, et al.
Veröffentlicht: (2025)
von: Zhao, Yushu, et al.
Veröffentlicht: (2025)
SparseDPD: A Sparse Neural Network-based Digital Predistortion FPGA Accelerator for RF Power Amplifier Linearization
von: Versluis, Manno, et al.
Veröffentlicht: (2025)
von: Versluis, Manno, et al.
Veröffentlicht: (2025)
WATOS: Efficient LLM Training Strategies and Architecture Co-exploration for Wafer-scale Chip
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
A High-Throughput Hardware Accelerator for Lempel-Ziv 4 Compression Algorithm
von: Chen, Tao, et al.
Veröffentlicht: (2024)
von: Chen, Tao, et al.
Veröffentlicht: (2024)
A Digital SRAM-Based Compute-In-Memory Macro for Weight-Stationary Dynamic Matrix Multiplication in Transformer Attention Score Computation
von: Yu, Jianyi, et al.
Veröffentlicht: (2025)
von: Yu, Jianyi, et al.
Veröffentlicht: (2025)
A Low-Latency FFT-IFFT Cascade Architecture
von: Parhi, Keshab K.
Veröffentlicht: (2023)
von: Parhi, Keshab K.
Veröffentlicht: (2023)
Pipeline Stage Resolved Timing Characterization of FPGA and ASIC Implementations of a RISC V Processor
von: Darvishi, Mostafa
Veröffentlicht: (2025)
von: Darvishi, Mostafa
Veröffentlicht: (2025)
Channel-Coherence-Adaptive Two-Stage Fully Digital Combining for mmWave MIMO Systems
von: Khorsandmanesh, Yasaman, et al.
Veröffentlicht: (2025)
von: Khorsandmanesh, Yasaman, et al.
Veröffentlicht: (2025)
RRAM-Based Bio-Inspired Circuits for Mobile Epileptic Correlation Extraction and Seizure Prediction
von: Wang, Hao, et al.
Veröffentlicht: (2024)
von: Wang, Hao, et al.
Veröffentlicht: (2024)
A Spatial Array for Spectrally Agile Wireless Processing
von: Rasteh, Ali, et al.
Veröffentlicht: (2025)
von: Rasteh, Ali, et al.
Veröffentlicht: (2025)
A Hierarchical Dataflow-Driven Heterogeneous Architecture for Wireless Baseband Processing
von: Jiang, Limin, et al.
Veröffentlicht: (2024)
von: Jiang, Limin, et al.
Veröffentlicht: (2024)
Lightator: An Optical Near-Sensor Accelerator with Compressive Acquisition Enabling Versatile Image Processing
von: Morsali, Mehrdad, et al.
Veröffentlicht: (2024)
von: Morsali, Mehrdad, et al.
Veröffentlicht: (2024)
Single-Event Upset Analysis of a Systolic Array based Deep Neural Network Accelerator
von: Jonckers, Naïn, et al.
Veröffentlicht: (2024)
von: Jonckers, Naïn, et al.
Veröffentlicht: (2024)
Design and In-training Optimization of Binary Search ADC for Flexible Classifiers
von: Duarte, Paula Carolina Lozano, et al.
Veröffentlicht: (2024)
von: Duarte, Paula Carolina Lozano, et al.
Veröffentlicht: (2024)
Practical Timing Closure in FPGA and ASIC Designs: Methods, Challenges, and Case Studies
von: Darvishi, Mostafa
Veröffentlicht: (2025)
von: Darvishi, Mostafa
Veröffentlicht: (2025)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Large Attention-Based Model Inference on Tile-Based Accelerators
von: Zhang, Chi, et al.
Veröffentlicht: (2026)
von: Zhang, Chi, et al.
Veröffentlicht: (2026)
STAR: An Efficient Softmax Engine for Attention Model with RRAM Crossbar
von: Zhai, Yifeng, et al.
Veröffentlicht: (2024)
von: Zhai, Yifeng, et al.
Veröffentlicht: (2024)
Decade-Bandwidth RF-Input Pseudo-Doherty Load Modulated Balanced Amplifier using Signal-Flow-Based Phase Alignment Design
von: Gong, Pingzhu, et al.
Veröffentlicht: (2024)
von: Gong, Pingzhu, et al.
Veröffentlicht: (2024)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
Accelerating Multi-Scale Deformable Attention Using Near-Memory-Processing Architecture
von: Li, Huize, et al.
Veröffentlicht: (2026)
von: Li, Huize, et al.
Veröffentlicht: (2026)
VIKIN: A Reconfigurable Accelerator for KANs and MLPs with Two-Stage Sparsity Support
von: Ou, Wenhui, et al.
Veröffentlicht: (2026)
von: Ou, Wenhui, et al.
Veröffentlicht: (2026)
Slimmed optical neural networks with multiplexed neuron sets and a corresponding backpropagation training algorithm
von: Liu, Yi-Feng, et al.
Veröffentlicht: (2023)
von: Liu, Yi-Feng, et al.
Veröffentlicht: (2023)
METRO: A Software-Hardware Co-Design of Interconnections for Spatial DNN Accelerators
von: Wang, Zhao, et al.
Veröffentlicht: (2021)
von: Wang, Zhao, et al.
Veröffentlicht: (2021)
The Quest for Reliable AI Accelerators: Cross-Layer Evaluation and Design Optimization
von: Li, Meng, et al.
Veröffentlicht: (2026)
von: Li, Meng, et al.
Veröffentlicht: (2026)
Monad: Towards Cost-effective Specialization for Chiplet-based Spatial Accelerators
von: Hao, Xiaochen, et al.
Veröffentlicht: (2023)
von: Hao, Xiaochen, et al.
Veröffentlicht: (2023)
Optimized Spatial Architecture Mapping Flow for Transformer Accelerators
von: Xu, Haocheng, et al.
Veröffentlicht: (2024)
von: Xu, Haocheng, et al.
Veröffentlicht: (2024)
GCC: A 3DGS Inference Architecture with Gaussian-Wise and Cross-Stage Conditional Processing
von: Pei, Minnan, et al.
Veröffentlicht: (2025)
von: Pei, Minnan, et al.
Veröffentlicht: (2025)
SnipSnap: A Joint Compression Format and Dataflow Co-Optimization Framework for Efficient Sparse LLM Accelerator Design
von: Wu, Junyi, et al.
Veröffentlicht: (2025)
von: Wu, Junyi, et al.
Veröffentlicht: (2025)
COBRA: Algorithm-Architecture Co-optimized Binary Transformer Accelerator for Edge Inference
von: Qiao, Ye, et al.
Veröffentlicht: (2025)
von: Qiao, Ye, et al.
Veröffentlicht: (2025)
SuperUROP: An FPGA-Based Spatial Accelerator for Sparse Matrix Operations
von: Parthasarathy, Rishab
Veröffentlicht: (2025)
von: Parthasarathy, Rishab
Veröffentlicht: (2025)
A Switch-Centric In-Network Architecture for Accelerating LLM Inference in Shared-Memory Network
von: Jiang, Aojie, et al.
Veröffentlicht: (2026)
von: Jiang, Aojie, et al.
Veröffentlicht: (2026)
A Lightweight Architecture for Real-Time Neuronal-Spike Classification
von: Siddiqi, Muhammad Ali, et al.
Veröffentlicht: (2023)
von: Siddiqi, Muhammad Ali, et al.
Veröffentlicht: (2023)
StreamDCIM: A Tile-based Streaming Digital CIM Accelerator with Mixed-stationary Cross-forwarding Dataflow for Multimodal Transformer
von: Qin, Shantian, et al.
Veröffentlicht: (2025)
von: Qin, Shantian, et al.
Veröffentlicht: (2025)
Efficient yet Accurate End-to-End SC Accelerator Design
von: Li, Meng, et al.
Veröffentlicht: (2024)
von: Li, Meng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
PADE: A Predictor-Free Sparse Attention Accelerator via Unified Execution and Stage Fusion
von: Wang, Huizheng, et al.
Veröffentlicht: (2025) -
MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness
von: Wang, Huizheng, et al.
Veröffentlicht: (2025) -
SOFA: A Compute-Memory Optimized Sparsity Accelerator via Cross-Stage Coordinated Tiling
von: Wang, Huizheng, et al.
Veröffentlicht: (2024) -
TEMP: A Memory Efficient Physical-aware Tensor Partition-Mapping Framework on Wafer-scale Chips
von: Wang, Huizheng, et al.
Veröffentlicht: (2025) -
Efficient Orchestrated AI Workflows Execution on Scale-out Spatial Architecture
von: Deng, Jinyi, et al.
Veröffentlicht: (2024)