SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Hanrui, Zhang, Zhekai, Han, Song |
|---|---|
| Format: | Preprint |
| Published: |
2020
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SpArch: Efficient Architecture for Sparse Matrix Multiplication
by: Zhang, Zhekai, et al.
Published: (2020)
by: Zhang, Zhekai, et al.
Published: (2020)
LEGO: Spatial Accelerator Generation and Optimization for Tensor Applications
by: Lin, Yujun, et al.
Published: (2025)
by: Lin, Yujun, et al.
Published: (2025)
Dynamic Sparse Attention: Access Patterns and Architecture
by: Levy, Noam
Published: (2026)
by: Levy, Noam
Published: (2026)
TurboAttention: Efficient Attention Approximation For High Throughputs LLMs
by: Kang, Hao, et al.
Published: (2024)
by: Kang, Hao, et al.
Published: (2024)
MonoSparse-CAM: Efficient Tree Model Processing via Monotonicity and Sparsity in CAMs
by: Molom-Ochir, Tergel, et al.
Published: (2024)
by: Molom-Ochir, Tergel, et al.
Published: (2024)
Enabling Unstructured Sparse Acceleration on Structured Sparse Accelerators
by: Jeong, Geonhwa, et al.
Published: (2024)
by: Jeong, Geonhwa, et al.
Published: (2024)
MEMHD: Memory-Efficient Multi-Centroid Hyperdimensional Computing for Fully-Utilized In-Memory Computing Architectures
by: Kang, Do Yeong, et al.
Published: (2025)
by: Kang, Do Yeong, et al.
Published: (2025)
tuGEMM: Area-Power-Efficient Temporal Unary GEMM Architecture for Low-Precision Edge AI
by: Nair, Harideep, et al.
Published: (2024)
by: Nair, Harideep, et al.
Published: (2024)
AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIM
by: Zhang, Yuanpeng, et al.
Published: (2025)
by: Zhang, Yuanpeng, et al.
Published: (2025)
Self-Attention to Operator Learning-based 3D-IC Thermal Simulation
by: Huang, Zhen, et al.
Published: (2025)
by: Huang, Zhen, et al.
Published: (2025)
Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inference
by: Chen, Hongzheng, et al.
Published: (2023)
by: Chen, Hongzheng, et al.
Published: (2023)
SymRTLO: Enhancing RTL Code Optimization with LLMs and Neuron-Inspired Symbolic Reasoning
by: Wang, Yiting, et al.
Published: (2025)
by: Wang, Yiting, et al.
Published: (2025)
QiMeng-CodeV-SVA: Training Specialized LLMs for Hardware Assertion Generation via RTL-Grounded Bidirectional Data Synthesis
by: Wu, Yutong, et al.
Published: (2026)
by: Wu, Yutong, et al.
Published: (2026)
The Graph's Apprentice: Teaching an LLM Low Level Knowledge for Circuit Quality Estimation
by: Moravej, Reza, et al.
Published: (2024)
by: Moravej, Reza, et al.
Published: (2024)
TRINE: A Token-Aware, Runtime-Adaptive FPGA Inference Engine for Multimodal AI
by: Oh, Hyunwoo, et al.
Published: (2026)
by: Oh, Hyunwoo, et al.
Published: (2026)
VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation
by: Wang, Yiting, et al.
Published: (2025)
by: Wang, Yiting, et al.
Published: (2025)
Lorecast: Layout-Aware Performance and Power Forecasting from Natural Language
by: Wang, Runzhi, et al.
Published: (2025)
by: Wang, Runzhi, et al.
Published: (2025)
MoNDE: Mixture of Near-Data Experts for Large-Scale Sparse Models
by: Kim, Taehyun, et al.
Published: (2024)
by: Kim, Taehyun, et al.
Published: (2024)
FLASH-D: FlashAttention with Hidden Softmax Division
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
Hardware/Software Co-Design of RISC-V Extensions for Accelerating Sparse DNNs on FPGAs
by: Sabih, Muhammad, et al.
Published: (2025)
by: Sabih, Muhammad, et al.
Published: (2025)
The Role of Advanced Computer Architectures in Accelerating Artificial Intelligence Workloads
by: Amin, Shahid, et al.
Published: (2025)
by: Amin, Shahid, et al.
Published: (2025)
MixPE: Quantization and Hardware Co-design for Efficient LLM Inference
by: Zhang, Yu, et al.
Published: (2024)
by: Zhang, Yu, et al.
Published: (2024)
FAMOUS: Flexible Accelerator for the Attention Mechanism of Transformer on UltraScale+ FPGAs
by: Kabir, Ehsan, et al.
Published: (2024)
by: Kabir, Ehsan, et al.
Published: (2024)
Revolutionizing TCAD Simulations with Universal Device Encoding and Graph Attention Networks
by: Fan, Guangxi, et al.
Published: (2023)
by: Fan, Guangxi, et al.
Published: (2023)
QuArch: A Question-Answering Dataset for AI Agents in Computer Architecture
by: Prakash, Shvetank, et al.
Published: (2025)
by: Prakash, Shvetank, et al.
Published: (2025)
From Fuzzy to Exact: The Halo Architecture for Infinite-Depth Reasoning via Rational Arithmetic
by: Ren, Hansheng
Published: (2026)
by: Ren, Hansheng
Published: (2026)
VUSA: Virtually Upscaled Systolic Array Architecture to Exploit Unstructured Sparsity in AI Acceleration
by: Helal, Shereef, et al.
Published: (2025)
by: Helal, Shereef, et al.
Published: (2025)
Multi-objective Optimization in CPU Design Space Exploration: Attention is All You Need
by: Xue, Runzhen, et al.
Published: (2024)
by: Xue, Runzhen, et al.
Published: (2024)
Mitigating hallucinations and omissions in LLMs for invertible problems: An application to hardware logic design automation
by: Cassidy, Andrew S., et al.
Published: (2025)
by: Cassidy, Andrew S., et al.
Published: (2025)
hdl2v: A Code Translation Dataset for Enhanced LLM Verilog Generation
by: Hong, Charles, et al.
Published: (2025)
by: Hong, Charles, et al.
Published: (2025)
VeriMind: Agentic LLM for Automated Verilog Generation with a Novel Evaluation Metric
by: Nadimi, Bardia, et al.
Published: (2025)
by: Nadimi, Bardia, et al.
Published: (2025)
Enabling New HDLs with Agents
by: Zakharov, Mark, et al.
Published: (2024)
by: Zakharov, Mark, et al.
Published: (2024)
Highly Optimized Kernels and Fine-Grained Codebooks for LLM Inference on Arm CPUs
by: Gope, Dibakar, et al.
Published: (2024)
by: Gope, Dibakar, et al.
Published: (2024)
Autocomp: A Powerful and Portable Code Optimizer for Tensor Accelerators
by: Hong, Charles, et al.
Published: (2025)
by: Hong, Charles, et al.
Published: (2025)
Chameleon: a Heterogeneous and Disaggregated Accelerator System for Retrieval-Augmented Language Models
by: Jiang, Wenqi, et al.
Published: (2023)
by: Jiang, Wenqi, et al.
Published: (2023)
PyraNet: A Multi-Layered Hierarchical Dataset for Verilog
by: Nadimi, Bardia, et al.
Published: (2024)
by: Nadimi, Bardia, et al.
Published: (2024)
Efficient Arbitrary Precision Acceleration for Large Language Models on GPU Tensor Cores
by: Ma, Shaobo, et al.
Published: (2024)
by: Ma, Shaobo, et al.
Published: (2024)
Zero-Shot RTL Code Generation with Attention Sink Augmented Large Language Models
by: Sandal, Selim, et al.
Published: (2024)
by: Sandal, Selim, et al.
Published: (2024)
FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design
by: Xia, Haojun, et al.
Published: (2024)
by: Xia, Haojun, et al.
Published: (2024)
AttentionLego: An Open-Source Building Block For Spatially-Scalable Large Language Model Accelerator With Processing-In-Memory Technology
by: Cong, Rongqing, et al.
Published: (2024)
by: Cong, Rongqing, et al.
Published: (2024)
Similar Items
-
SpArch: Efficient Architecture for Sparse Matrix Multiplication
by: Zhang, Zhekai, et al.
Published: (2020) -
LEGO: Spatial Accelerator Generation and Optimization for Tensor Applications
by: Lin, Yujun, et al.
Published: (2025) -
Dynamic Sparse Attention: Access Patterns and Architecture
by: Levy, Noam
Published: (2026) -
TurboAttention: Efficient Attention Approximation For High Throughputs LLMs
by: Kang, Hao, et al.
Published: (2024) -
MonoSparse-CAM: Efficient Tree Model Processing via Monotonicity and Sparsity in CAMs
by: Molom-Ochir, Tergel, et al.
Published: (2024)