AttentionEngine: A Versatile Framework for Efficient Attention Mechanisms on Diverse Hardware Platforms
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Feiyang, Cheng, Yu, Wang, Lei, Xia, Yuqing, Miao, Ziming, Ma, Lingxiao, Yang, Fan, Xue, Jilong, Yang, Zhi, Yang, Mao, Chen, Haibo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DRIM-ANN: An Approximate Nearest Neighbor Search Engine based on Commercial DRAM-PIMs
por: Chen, Mingkai, et al.
Publicado: (2024)
por: Chen, Mingkai, et al.
Publicado: (2024)
Long-term Monitoring of Kernel and Hardware Events to Understand Latency Variance
por: Zhou, Fang, et al.
Publicado: (2026)
por: Zhou, Fang, et al.
Publicado: (2026)
FlashOmni: A Unified Sparse Attention Engine for Diffusion Transformers
por: Qiao, Liang, et al.
Publicado: (2025)
por: Qiao, Liang, et al.
Publicado: (2025)
Cloud Computing Energy Consumption Prediction Based on Kernel Extreme Learning Machine Algorithm Improved by Vector Weighted Average Algorithm
por: Wang, Yuqing, et al.
Publicado: (2025)
por: Wang, Yuqing, et al.
Publicado: (2025)
Towards Efficient Multi-Scale Deformable Attention on NPU
por: Huang, Chenghuan, et al.
Publicado: (2025)
por: Huang, Chenghuan, et al.
Publicado: (2025)
H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference
por: Fu, Zizhuo, et al.
Publicado: (2025)
por: Fu, Zizhuo, et al.
Publicado: (2025)
Efficient Hardware Accelerator Based on Medium Granularity Dataflow for SpTRSV
por: Chen, Qian, et al.
Publicado: (2024)
por: Chen, Qian, et al.
Publicado: (2024)
Dissecting RISC-V Performance: Practical PMU Profiling and Hardware-Agnostic Roofline Analysis on Emerging Platforms
por: Batashev, Alexander
Publicado: (2025)
por: Batashev, Alexander
Publicado: (2025)
Atys: An Efficient Profiling Framework for Identifying Hotspot Functions in Large-scale Cloud Microservices
por: Sun, Jiaqi, et al.
Publicado: (2025)
por: Sun, Jiaqi, et al.
Publicado: (2025)
Scaler: Efficient and Effective Cross Flow Analysis
por: Steven, et al.
Publicado: (2024)
por: Steven, et al.
Publicado: (2024)
DF-GNN: Dynamic Fusion Framework for Attention Graph Neural Networks on GPUs
por: Liu, Jiahui, et al.
Publicado: (2024)
por: Liu, Jiahui, et al.
Publicado: (2024)
CAPSim: A Fast CPU Performance Simulator Using Attention-based Predictor
por: Xu, Buqing, et al.
Publicado: (2025)
por: Xu, Buqing, et al.
Publicado: (2025)
Matryoshka: Optimization of Dynamic Diverse Quantum Chemistry Systems via Elastic Parallelism Transformation
por: Wang, Tuowei, et al.
Publicado: (2024)
por: Wang, Tuowei, et al.
Publicado: (2024)
Flex Attention: A Programming Model for Generating Optimized Attention Kernels
por: Dong, Juechu, et al.
Publicado: (2024)
por: Dong, Juechu, et al.
Publicado: (2024)
Efficient Graph Knowledge Distillation from GNNs to Kolmogorov--Arnold Networks via Self-Attention Dynamic Sampling
por: Cui, Can, et al.
Publicado: (2025)
por: Cui, Can, et al.
Publicado: (2025)
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference
por: Yao, Feiyu, et al.
Publicado: (2026)
por: Yao, Feiyu, et al.
Publicado: (2026)
SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization
por: Zhang, Jintao, et al.
Publicado: (2024)
por: Zhang, Jintao, et al.
Publicado: (2024)
Design and implementation of a distributed security threat detection system integrating federated learning and multimodal LLM
por: Wang, Yuqing, et al.
Publicado: (2025)
por: Wang, Yuqing, et al.
Publicado: (2025)
Block Sparse Flash Attention
por: Ohayon, Daniel, et al.
Publicado: (2025)
por: Ohayon, Daniel, et al.
Publicado: (2025)
Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants
por: You, Bozhi, et al.
Publicado: (2025)
por: You, Bozhi, et al.
Publicado: (2025)
LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
por: Yang, Shang, et al.
Publicado: (2025)
por: Yang, Shang, et al.
Publicado: (2025)
Achieving Consistent and Comparable CPU Evaluation
por: Wang, Chenxi, et al.
Publicado: (2024)
por: Wang, Chenxi, et al.
Publicado: (2024)
SAfEPaTh: A System-Level Approach for Efficient Power and Thermal Estimation of Convolutional Neural Network Accelerator
por: Chen, Yukai, et al.
Publicado: (2024)
por: Chen, Yukai, et al.
Publicado: (2024)
SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference
por: Zhang, Jintao, et al.
Publicado: (2025)
por: Zhang, Jintao, et al.
Publicado: (2025)
Multi-Strided Access Patterns to Boost Hardware Prefetching
por: Blom, Miguel O., et al.
Publicado: (2024)
por: Blom, Miguel O., et al.
Publicado: (2024)
CXL and the Return of Scale-Up Database Engines
por: Lerner, Alberto, et al.
Publicado: (2024)
por: Lerner, Alberto, et al.
Publicado: (2024)
Ragged Paged Attention: A High-Performance and Flexible LLM Inference Kernel for TPU
por: Jiang, Jevin, et al.
Publicado: (2026)
por: Jiang, Jevin, et al.
Publicado: (2026)
Spatiotemporal Non-Uniformity-Aware Online Task Scheduling in Collaborative Edge Computing for Industrial Internet of Things
por: Li, Yang, et al.
Publicado: (2025)
por: Li, Yang, et al.
Publicado: (2025)
Attention in SRAM on Tenstorrent Grayskull
por: Thüning, Moritz
Publicado: (2024)
por: Thüning, Moritz
Publicado: (2024)
DyLLM: Efficient Diffusion LLM Inference via Saliency-based Token Selection and Partial Attention
por: Lee, Younjoo, et al.
Publicado: (2026)
por: Lee, Younjoo, et al.
Publicado: (2026)
A Latency-Constrained, Gated Recurrent Unit (GRU) Implementation in the Versal AI Engine
por: Sapkas, M., et al.
Publicado: (2025)
por: Sapkas, M., et al.
Publicado: (2025)
Dynamic Precision Math Engine for Linear Algebra and Trigonometry Acceleration on Xtensa LX6 Microcontrollers
por: Preciado, Elian Alfonso Lopez
Publicado: (2026)
por: Preciado, Elian Alfonso Lopez
Publicado: (2026)
Sawtooth Wavefront Reordering: Enhanced CuTile FlashAttention on NVIDIA GB10
por: Zhu, Yifan, et al.
Publicado: (2026)
por: Zhu, Yifan, et al.
Publicado: (2026)
Karatsuba Matrix Multiplication and its Efficient Custom Hardware Implementations
por: Pogue, Trevor E., et al.
Publicado: (2025)
por: Pogue, Trevor E., et al.
Publicado: (2025)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
por: Zhao, Xuanlei, et al.
Publicado: (2024)
por: Zhao, Xuanlei, et al.
Publicado: (2024)
AutoChunk: Automated Activation Chunk for Memory-Efficient Long Sequence Inference
por: Zhao, Xuanlei, et al.
Publicado: (2024)
por: Zhao, Xuanlei, et al.
Publicado: (2024)
Uncertainty Quantification as a Complementary Latent Health Indicator for Remaining Useful Life Prediction on Turbofan Engines
por: Thil, Lucas, et al.
Publicado: (2025)
por: Thil, Lucas, et al.
Publicado: (2025)
WaveTune: Wave-aware Bilinear Modeling for Efficient GPU Kernel Auto-tuning
por: Zhang, Kaixuan, et al.
Publicado: (2026)
por: Zhang, Kaixuan, et al.
Publicado: (2026)
Online Pseudo-average Shifting Attention(PASA) for Robust Low-precision LLM Inference: Algorithms and Numerical Analysis
por: Cheng, Long, et al.
Publicado: (2025)
por: Cheng, Long, et al.
Publicado: (2025)
KForge: Program Synthesis for Diverse AI Hardware Accelerators
por: Sereda, Taras, et al.
Publicado: (2025)
por: Sereda, Taras, et al.
Publicado: (2025)
Ejemplares similares
-
DRIM-ANN: An Approximate Nearest Neighbor Search Engine based on Commercial DRAM-PIMs
por: Chen, Mingkai, et al.
Publicado: (2024) -
Long-term Monitoring of Kernel and Hardware Events to Understand Latency Variance
por: Zhou, Fang, et al.
Publicado: (2026) -
FlashOmni: A Unified Sparse Attention Engine for Diffusion Transformers
por: Qiao, Liang, et al.
Publicado: (2025) -
Cloud Computing Energy Consumption Prediction Based on Kernel Extreme Learning Machine Algorithm Improved by Vector Weighted Average Algorithm
por: Wang, Yuqing, et al.
Publicado: (2025) -
Towards Efficient Multi-Scale Deformable Attention on NPU
por: Huang, Chenghuan, et al.
Publicado: (2025)