TurboAttention: Efficient Attention Approximation For High Throughputs LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Kang, Hao, Bharadwaj, Srikant, Hensman, James, Krishna, Tushar, Ruhle, Victor, Rajmohan, Saravan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lean Attention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers
by: Sanovar, Rya, et al.
Published: (2024)
by: Sanovar, Rya, et al.
Published: (2024)
Taming the Tail: NoI Topology Synthesis for Mixed DL Workloads on Chiplet-Based Accelerators
by: Shukla, Arnav, et al.
Published: (2025)
by: Shukla, Arnav, et al.
Published: (2025)
MicroScopiQ: Accelerating Foundational Models through Outlier-Aware Microscaling Quantization
by: Ramachandran, Akshat, et al.
Published: (2024)
by: Ramachandran, Akshat, et al.
Published: (2024)
Design Conductor 2.0: An agent builds a TurboQuant inference accelerator in 80 hours
by: The Verkor Team, et al.
Published: (2026)
by: The Verkor Team, et al.
Published: (2026)
AIRCHITECT v2: Learning the Hardware Accelerator Design Space through Unified Representations
by: Seo, Jamin, et al.
Published: (2025)
by: Seo, Jamin, et al.
Published: (2025)
PIM-LLM: A High-Throughput Hybrid PIM Architecture for 1-bit LLMs
by: Malekar, Jinendra, et al.
Published: (2025)
by: Malekar, Jinendra, et al.
Published: (2025)
Multi-objective Optimization in CPU Design Space Exploration: Attention is All You Need
by: Xue, Runzhen, et al.
Published: (2024)
by: Xue, Runzhen, et al.
Published: (2024)
Enabling Unstructured Sparse Acceleration on Structured Sparse Accelerators
by: Jeong, Geonhwa, et al.
Published: (2024)
by: Jeong, Geonhwa, et al.
Published: (2024)
Dynamic Sparse Attention: Access Patterns and Architecture
by: Levy, Noam
Published: (2026)
by: Levy, Noam
Published: (2026)
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
by: Ramachandran, Akshat, et al.
Published: (2025)
by: Ramachandran, Akshat, et al.
Published: (2025)
FLASH-D: FlashAttention with Hidden Softmax Division
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
SWAT: Scalable and Efficient Window Attention-based Transformers Acceleration on FPGAs
by: Bai, Zhenyu, et al.
Published: (2024)
by: Bai, Zhenyu, et al.
Published: (2024)
FAMOUS: Flexible Accelerator for the Attention Mechanism of Transformer on UltraScale+ FPGAs
by: Kabir, Ehsan, et al.
Published: (2024)
by: Kabir, Ehsan, et al.
Published: (2024)
Revolutionizing TCAD Simulations with Universal Device Encoding and Graph Attention Networks
by: Fan, Guangxi, et al.
Published: (2023)
by: Fan, Guangxi, et al.
Published: (2023)
SystolicAttention: Fusing FlashAttention within a Single Systolic Array
by: Lin, Jiawei, et al.
Published: (2025)
by: Lin, Jiawei, et al.
Published: (2025)
TReX- Reusing Vision Transformer's Attention for Efficient Xbar-based Computing
by: Moitra, Abhishek, et al.
Published: (2024)
by: Moitra, Abhishek, et al.
Published: (2024)
Self-Attention to Operator Learning-based 3D-IC Thermal Simulation
by: Huang, Zhen, et al.
Published: (2025)
by: Huang, Zhen, et al.
Published: (2025)
Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding
by: Fan, Wang, et al.
Published: (2026)
by: Fan, Wang, et al.
Published: (2026)
Hardware Efficient Approximate Convolution with Tunable Error Tolerance for CNNs
by: Shashidhar, Vishal, et al.
Published: (2026)
by: Shashidhar, Vishal, et al.
Published: (2026)
Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns
by: Bambhaniya, Abhimanyu, et al.
Published: (2026)
by: Bambhaniya, Abhimanyu, et al.
Published: (2026)
Algorithm-Hardware Co-Design of Distribution-Aware Logarithmic-Posit Encodings for Efficient DNN Inference
by: Ramachandran, Akshat, et al.
Published: (2024)
by: Ramachandran, Akshat, et al.
Published: (2024)
ARCS: Autoregressive Circuit Synthesis with Topology-Aware Graph Attention and Spec Conditioning
by: Pathak, Tushar Dhananjay
Published: (2026)
by: Pathak, Tushar Dhananjay
Published: (2026)
Low Power Approximate Multiplier Architecture for Deep Neural Networks
by: Jaswal, Pragun, et al.
Published: (2025)
by: Jaswal, Pragun, et al.
Published: (2025)
SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
by: Wang, Hanrui, et al.
Published: (2020)
by: Wang, Hanrui, et al.
Published: (2020)
Explainable AI-Guided Efficient Approximate DNN Generation for Multi-Pod Systolic Arrays
by: Siddique, Ayesha, et al.
Published: (2025)
by: Siddique, Ayesha, et al.
Published: (2025)
AttentionLego: An Open-Source Building Block For Spatially-Scalable Large Language Model Accelerator With Processing-In-Memory Technology
by: Cong, Rongqing, et al.
Published: (2024)
by: Cong, Rongqing, et al.
Published: (2024)
REASON: Accelerating Probabilistic Logical Reasoning for Scalable Neuro-Symbolic Intelligence
by: Wan, Zishen, et al.
Published: (2026)
by: Wan, Zishen, et al.
Published: (2026)
DPUV4E: High-Throughput DPU Architecture Design for CNN on Versal ACAP
by: Li, Guoyu, et al.
Published: (2025)
by: Li, Guoyu, et al.
Published: (2025)
Chiplet-Based RISC-V SoC with Modular AI Acceleration
by: Bharadwaj, Suhas Suresh, et al.
Published: (2025)
by: Bharadwaj, Suhas Suresh, et al.
Published: (2025)
SageAttention2++: A More Efficient Implementation of SageAttention2
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
NSFlow: An End-to-End FPGA Framework with Scalable Dataflow Architecture for Neuro-Symbolic AI
by: Yang, Hanchen, et al.
Published: (2025)
by: Yang, Hanchen, et al.
Published: (2025)
Reducing the Cost of Dropout in Flash-Attention by Hiding RNG with GEMM
by: Ma, Haiyue, et al.
Published: (2024)
by: Ma, Haiyue, et al.
Published: (2024)
HLS-Eval: A Benchmark and Framework for Evaluating LLMs on High-Level Synthesis Design Tasks
by: Abi-Karam, Stefan, et al.
Published: (2025)
by: Abi-Karam, Stefan, et al.
Published: (2025)
MEMHD: Memory-Efficient Multi-Centroid Hyperdimensional Computing for Fully-Utilized In-Memory Computing Architectures
by: Kang, Do Yeong, et al.
Published: (2025)
by: Kang, Do Yeong, et al.
Published: (2025)
Rethinking LLM Inference Bottlenecks: Insights from Latent Attention and Mixture-of-Experts
by: Yun, Sungmin, et al.
Published: (2025)
by: Yun, Sungmin, et al.
Published: (2025)
Static IR Drop Prediction with Attention U-Net and Saliency-Based Explainability
by: Zhang, Lizi, et al.
Published: (2024)
by: Zhang, Lizi, et al.
Published: (2024)
Column-wise Quantization of Weights and Partial Sums for Accurate and Efficient Compute-In-Memory Accelerators
by: Kim, Jiyoon, et al.
Published: (2025)
by: Kim, Jiyoon, et al.
Published: (2025)
GTAC: A Generative Transformer for Approximate Circuits
by: Wang, Jingxin, et al.
Published: (2025)
by: Wang, Jingxin, et al.
Published: (2025)
MAGNet: A Multi-Scale Attention-Guided Graph Fusion Network for DRC Violation Detection
by: Lu, Weihan, et al.
Published: (2025)
by: Lu, Weihan, et al.
Published: (2025)
HLSTransform: Energy-Efficient Llama 2 Inference on FPGAs Via High Level Synthesis
by: He, Andy, et al.
Published: (2024)
by: He, Andy, et al.
Published: (2024)
Similar Items
-
Lean Attention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers
by: Sanovar, Rya, et al.
Published: (2024) -
Taming the Tail: NoI Topology Synthesis for Mixed DL Workloads on Chiplet-Based Accelerators
by: Shukla, Arnav, et al.
Published: (2025) -
MicroScopiQ: Accelerating Foundational Models through Outlier-Aware Microscaling Quantization
by: Ramachandran, Akshat, et al.
Published: (2024) -
Design Conductor 2.0: An agent builds a TurboQuant inference accelerator in 80 hours
by: The Verkor Team, et al.
Published: (2026) -
AIRCHITECT v2: Learning the Hardware Accelerator Design Space through Unified Representations
by: Seo, Jamin, et al.
Published: (2025)