Lean Attention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Sanovar, Rya, Bharadwaj, Srikant, Amant, Renee St., Rühle, Victor, Rajmohan, Saravan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TurboAttention: Efficient Attention Approximation For High Throughputs LLMs
by: Kang, Hao, et al.
Published: (2024)
by: Kang, Hao, et al.
Published: (2024)
LIPPEN: A Lightweight In-Place Pointer Encryption Architecture for Pointer Integrity
by: Iravani, Erfan, et al.
Published: (2026)
by: Iravani, Erfan, et al.
Published: (2026)
Not All Thoughts Need HBM: Semantics-Aware Memory Hierarchy for LLM Reasoning
by: Yuan, Aojie, et al.
Published: (2026)
by: Yuan, Aojie, et al.
Published: (2026)
TokenStack: A Heterogeneous HBM-PIM Architecture and Runtime for Efficient LLM Inference
by: Li, Zhuoran, et al.
Published: (2026)
by: Li, Zhuoran, et al.
Published: (2026)
FOCUS: DLLMs Know How to Tame Their Compute Bound
by: Liang, Kaihua, et al.
Published: (2026)
by: Liang, Kaihua, et al.
Published: (2026)
REIS: A High-Performance and Energy-Efficient Retrieval System with In-Storage Processing
by: Chen, Kangqi, et al.
Published: (2025)
by: Chen, Kangqi, et al.
Published: (2025)
Architecture-Aware LLM Inference Optimization on AMD Instinct GPUs: A Comprehensive Benchmark and Deployment Study
by: Georgiou, Athos
Published: (2026)
by: Georgiou, Athos
Published: (2026)
Exploring LLM-based Verilog Code Generation with Data-Efficient Fine-Tuning and Testbench Automation
by: Chen, Mu-Chi, et al.
Published: (2026)
by: Chen, Mu-Chi, et al.
Published: (2026)
Systolic Arrays and Structured Pruning Co-design for Efficient Transformers in Edge Systems
by: Palacios, Pedro, et al.
Published: (2024)
by: Palacios, Pedro, et al.
Published: (2024)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
by: Jo, Myeong Jun
Published: (2026)
by: Jo, Myeong Jun
Published: (2026)
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models
by: Couturier, Camille, et al.
Published: (2025)
by: Couturier, Camille, et al.
Published: (2025)
SiliconMind-V1: Multi-Agent Distillation and Debug-Reasoning Workflows for Verilog Code Generation
by: Chen, Mu-Chi, et al.
Published: (2026)
by: Chen, Mu-Chi, et al.
Published: (2026)
RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry
by: Lv, Bo, et al.
Published: (2026)
by: Lv, Bo, et al.
Published: (2026)
Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding
by: Fan, Wang, et al.
Published: (2026)
by: Fan, Wang, et al.
Published: (2026)
Nanoscaling Floating-Point (NxFP): NanoMantissa, Adaptive Microexponents, and Code Recycling for Direct-Cast Compression of Large Language Models
by: Lo, Yun-Chen, et al.
Published: (2024)
by: Lo, Yun-Chen, et al.
Published: (2024)
Biological Intuition on Digital Hardware: An RTL Implementation of Poisson-Encoded SNNs for Static Image Classification
by: Das, Debabrata, et al.
Published: (2026)
by: Das, Debabrata, et al.
Published: (2026)
Keyformer: KV Cache Reduction through Key Tokens Selection for Efficient Generative Inference
by: Adnan, Muhammad, et al.
Published: (2024)
by: Adnan, Muhammad, et al.
Published: (2024)
CPU Simulation Using Two-Phase Stratified Sampling
by: Ekman, Magnus
Published: (2026)
by: Ekman, Magnus
Published: (2026)
RayFlex: An Open-Source RTL Implementation of the Hardware Ray Tracer Datapath
by: Shen, Fangjia, et al.
Published: (2024)
by: Shen, Fangjia, et al.
Published: (2024)
Nonvolatile Charge-Domain Attention with HZO Ferroelectric Capacitors: A Simulation-Based Device-to-System Evaluation
by: Abouagour, Faris
Published: (2026)
by: Abouagour, Faris
Published: (2026)
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
by: Ganjihal, Sanjeev Rao
Published: (2026)
by: Ganjihal, Sanjeev Rao
Published: (2026)
Hardware-Centric Analysis of DeepSeek's Multi-Head Latent Attention
by: Geens, Robin, et al.
Published: (2025)
by: Geens, Robin, et al.
Published: (2025)
Hybrid Photonic-digital Accelerator for Attention Mechanism
by: Li, Huize, et al.
Published: (2025)
by: Li, Huize, et al.
Published: (2025)
SR-NCL: an Area-/Energy-Efficient Resilient NCL Architecture Based on Selective Redundancy
by: Ziad, Hasnain A., et al.
Published: (2025)
by: Ziad, Hasnain A., et al.
Published: (2025)
A flexible framework for early power and timing comparison of time-multiplexed CGRA kernel executions
by: Aspros, Maxime Henri, et al.
Published: (2025)
by: Aspros, Maxime Henri, et al.
Published: (2025)
SWAT: Scalable and Efficient Window Attention-based Transformers Acceleration on FPGAs
by: Bai, Zhenyu, et al.
Published: (2024)
by: Bai, Zhenyu, et al.
Published: (2024)
H-FA: A Hybrid Floating-Point and Logarithmic Approach to Hardware Accelerated FlashAttention
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
SATA: Sparsity-Aware Scheduling for Selective Token Attention
by: Fan, Zhenkun, et al.
Published: (2026)
by: Fan, Zhenkun, et al.
Published: (2026)
A Hardware-Based Multi-Stage Dynamic Power Management Architecture for Autonomous Low-Light Operation
by: Kouzinopoulos, Charalampos S., et al.
Published: (2026)
by: Kouzinopoulos, Charalampos S., et al.
Published: (2026)
QED: Scalable Verification of Hardware Memory Consistency
by: Ravi, Gokulan, et al.
Published: (2024)
by: Ravi, Gokulan, et al.
Published: (2024)
Photonic AI: A Hybrid Diffractive Holographic Neural System for Passive Optical Real-Time Image Classification
by: Hiremath, Prakul Sunil
Published: (2026)
by: Hiremath, Prakul Sunil
Published: (2026)
OpenEye: A Scalable Open-Source Hardware Accelerator for DNNs
by: Lebold, Denis, et al.
Published: (2026)
by: Lebold, Denis, et al.
Published: (2026)
Accelerating OTA Circuit Design: Transistor Sizing Based on a Transformer Model and Precomputed Lookup Tables
by: Ghosh, Subhadip, et al.
Published: (2025)
by: Ghosh, Subhadip, et al.
Published: (2025)
Hierarchical Shift Mixing -- Beyond Dense Attention in Transformers
by: Forchheimer, Robert
Published: (2026)
by: Forchheimer, Robert
Published: (2026)
CPU Simulation with Ranked Set Sampling and Repeated Subsampling
by: Ekman, Magnus
Published: (2026)
by: Ekman, Magnus
Published: (2026)
SwiftKV: An Edge-Oriented Attention Algorithm and Multi-Head Accelerator for Fast, Efficient LLM Decoding
by: Zhang, Junming, et al.
Published: (2026)
by: Zhang, Junming, et al.
Published: (2026)
Minimal Neuron Circuits -- Part I: Resonators
by: Nabil, Amr, et al.
Published: (2025)
by: Nabil, Amr, et al.
Published: (2025)
Minimal Neuron Circuits: Bursters
by: Nabil, Amr, et al.
Published: (2025)
by: Nabil, Amr, et al.
Published: (2025)
RISC-V Based TinyML Accelerator for Depthwise Separable Convolutions in Edge AI
by: Yildirim, Muhammed, et al.
Published: (2025)
by: Yildirim, Muhammed, et al.
Published: (2025)
Multiplier-free In-Memory Vector-Matrix Multiplication Using Distributed Arithmetic
by: Zeller, Felix, et al.
Published: (2025)
by: Zeller, Felix, et al.
Published: (2025)
Similar Items
-
TurboAttention: Efficient Attention Approximation For High Throughputs LLMs
by: Kang, Hao, et al.
Published: (2024) -
LIPPEN: A Lightweight In-Place Pointer Encryption Architecture for Pointer Integrity
by: Iravani, Erfan, et al.
Published: (2026) -
Not All Thoughts Need HBM: Semantics-Aware Memory Hierarchy for LLM Reasoning
by: Yuan, Aojie, et al.
Published: (2026) -
TokenStack: A Heterogeneous HBM-PIM Architecture and Runtime for Efficient LLM Inference
by: Li, Zhuoran, et al.
Published: (2026) -
FOCUS: DLLMs Know How to Tame Their Compute Bound
by: Liang, Kaihua, et al.
Published: (2026)