Zen-Attention: A Compiler Framework for Dynamic Attention Folding on AMD NPUs
Fuente:
arXiv
Guardado en:
| Autores principales: | Deshmukh, Aadesh, Raparti, Venkata Yaswanth, Hsu, Samuel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Optimizing sDTW for AMD GPUs
por: Latta-Lin, Daniel, et al.
Publicado: (2024)
por: Latta-Lin, Daniel, et al.
Publicado: (2024)
SGEMM-cube: Precision-Recovery FP32 GEMM Approximation on Ascend NPUs with FP16 Matrix Engines
por: Xue, Weicheng, et al.
Publicado: (2025)
por: Xue, Weicheng, et al.
Publicado: (2025)
LAPIS: A Performance Portable, High Productivity Compiler Framework
por: Kelley, Brian, et al.
Publicado: (2025)
por: Kelley, Brian, et al.
Publicado: (2025)
W4A16 Mixed-Precision Matrix Multiplication on Decoupled Architecture: Kernel Design and Memory Bottleneck Analysis for Ascend NPUs
por: He, Yuanhong, et al.
Publicado: (2026)
por: He, Yuanhong, et al.
Publicado: (2026)
A-IO: Adaptive Inference Orchestration for Memory-Bound NPUs
por: Zhang, Chen, et al.
Publicado: (2026)
por: Zhang, Chen, et al.
Publicado: (2026)
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
por: Zhou, Qihui, et al.
Publicado: (2025)
por: Zhou, Qihui, et al.
Publicado: (2025)
Understanding Data Movement in AMD Multi-GPU Systems with Infinity Fabric
por: Schieffer, Gabin, et al.
Publicado: (2024)
por: Schieffer, Gabin, et al.
Publicado: (2024)
exa-AMD: A Scalable Workflow for Accelerating AI-Assisted Materials Discovery and Design
por: Moraru, Maxim, et al.
Publicado: (2025)
por: Moraru, Maxim, et al.
Publicado: (2025)
BurstAttention: An Efficient Distributed Attention Framework for Extremely Long Sequences
por: Sun, Ao, et al.
Publicado: (2024)
por: Sun, Ao, et al.
Publicado: (2024)
Performance Portable Monte Carlo Particle Transport on Intel, NVIDIA, and AMD GPUs
por: Tramm, John, et al.
Publicado: (2024)
por: Tramm, John, et al.
Publicado: (2024)
DeepCompile: A Compiler-Driven Approach to Optimizing Distributed Deep Learning Training
por: Tanaka, Masahiro, et al.
Publicado: (2025)
por: Tanaka, Masahiro, et al.
Publicado: (2025)
From Attention to Disaggregation: Tracing the Evolution of LLM Inference
por: Kumar, Madabattula Rajesh, et al.
Publicado: (2025)
por: Kumar, Madabattula Rajesh, et al.
Publicado: (2025)
Inter-APU Communication on AMD MI300A Systems via Infinity Fabric: a Deep Dive
por: Schieffer, Gabin, et al.
Publicado: (2025)
por: Schieffer, Gabin, et al.
Publicado: (2025)
MT4G: A Tool for Reliable Auto-Discovery of NVIDIA and AMD GPU Compute and Memory Topologies
por: Vanecek, Stepan, et al.
Publicado: (2025)
por: Vanecek, Stepan, et al.
Publicado: (2025)
Mapping Parallel Matrix Multiplication in GotoBLAS2 to the AMD Versal ACAP for Deep Learning
por: Lei, Jie, et al.
Publicado: (2024)
por: Lei, Jie, et al.
Publicado: (2024)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
por: Zhang, Zhexiang, et al.
Publicado: (2025)
por: Zhang, Zhexiang, et al.
Publicado: (2025)
Porting HPC Applications to AMD Instinct$^\text{TM}$ MI300A Using Unified Memory and OpenMP
por: Tandon, Suyash, et al.
Publicado: (2024)
por: Tandon, Suyash, et al.
Publicado: (2024)
Is Flash Attention Stable?
por: Golden, Alicia, et al.
Publicado: (2024)
por: Golden, Alicia, et al.
Publicado: (2024)
Opt-GPTQ: An Optimized GPTQ Combining Sparse Attention and Quantization Techniques
por: Kong, Jie, et al.
Publicado: (2025)
por: Kong, Jie, et al.
Publicado: (2025)
Hestia: Hyperthread-Level Scheduling for Cloud Microservices with Interference-Aware Attention
por: Yang, Dingyu, et al.
Publicado: (2026)
por: Yang, Dingyu, et al.
Publicado: (2026)
Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference
por: Deshmukh, Dhruv, et al.
Publicado: (2025)
por: Deshmukh, Dhruv, et al.
Publicado: (2025)
Revealing the Challenges of Attention-FFN Disaggregation for Modern MoE Models and Hardware Systems
por: Liu, Guowei, et al.
Publicado: (2026)
por: Liu, Guowei, et al.
Publicado: (2026)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
por: Mo, Zizhao, et al.
Publicado: (2026)
por: Mo, Zizhao, et al.
Publicado: (2026)
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
por: Wahlgren, Jacob, et al.
Publicado: (2025)
por: Wahlgren, Jacob, et al.
Publicado: (2025)
AB-Sparse: Sparse Attention with Adaptive Block Size for Accurate and Efficient Long-Context Inference
por: Liu, Di, et al.
Publicado: (2026)
por: Liu, Di, et al.
Publicado: (2026)
S-HPLB: Efficient LLM Attention Serving via Sparsity-Aware Head Parallelism Load Balance
por: Liu, Di, et al.
Publicado: (2026)
por: Liu, Di, et al.
Publicado: (2026)
Preliminary report: Initial evaluation of StdPar implementations on AMD GPUs for HPC
por: Lin, Wei-Chen, et al.
Publicado: (2024)
por: Lin, Wei-Chen, et al.
Publicado: (2024)
Gensor: A Graph-based Construction Tensor Compilation Method for Deep Learning
por: Liu, Hangda, et al.
Publicado: (2025)
por: Liu, Hangda, et al.
Publicado: (2025)
Ksurf: Attention Kalman Filter and Principal Component Analysis for Prediction under Highly Variable Cloud Workloads
por: Dang'ana, Michael, et al.
Publicado: (2024)
por: Dang'ana, Michael, et al.
Publicado: (2024)
Efficient Parallel Compilation and Profiling of Quantum Circuits at Large Scales
por: Moore, Jane, et al.
Publicado: (2026)
por: Moore, Jane, et al.
Publicado: (2026)
Advancing Blockchain Scalability: A Linear Optimization Framework for Diversified Node Allocation in Shards
por: Assmann, Björn, et al.
Publicado: (2024)
por: Assmann, Björn, et al.
Publicado: (2024)
Mesh-Attention: A New Communication-Efficient Distributed Attention with Improved Data Locality
por: Chen, Sirui, et al.
Publicado: (2025)
por: Chen, Sirui, et al.
Publicado: (2025)
EAT: QoS-Aware Edge-Collaborative AIGC Task Scheduling via Attention-Guided Diffusion Reinforcement Learning
por: Xu, Zhifei, et al.
Publicado: (2025)
por: Xu, Zhifei, et al.
Publicado: (2025)
Bringing Auto-tuning to HIP: Analysis of Tuning Impact and Difficulty on AMD and Nvidia GPUs
por: Lurati, Milo, et al.
Publicado: (2024)
por: Lurati, Milo, et al.
Publicado: (2024)
Mitigating Temporal Blindness in Kubernetes Autoscaling: An Attention-Double-LSTM Framework
por: Shaikh, Faraz, et al.
Publicado: (2026)
por: Shaikh, Faraz, et al.
Publicado: (2026)
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
por: Zheng, Size, et al.
Publicado: (2025)
por: Zheng, Size, et al.
Publicado: (2025)
Flint: Compiler Enabled Cluster-Free Design Space Exploration for Distributed ML
por: Yoo, Jinsun, et al.
Publicado: (2026)
por: Yoo, Jinsun, et al.
Publicado: (2026)
Developing a BLAS library for the AMD AI Engine
por: Laan, Tristan, et al.
Publicado: (2024)
por: Laan, Tristan, et al.
Publicado: (2024)
iDynamics: A Configurable Emulation Framework for Evaluating Microservice Scheduling Policies under Controllable Cloud-Edge Dynamics
por: Chen, Ming, et al.
Publicado: (2025)
por: Chen, Ming, et al.
Publicado: (2025)
Joint Training on AMD and NVIDIA GPUs
por: Hu, Jon, et al.
Publicado: (2026)
por: Hu, Jon, et al.
Publicado: (2026)
Ejemplares similares
-
Optimizing sDTW for AMD GPUs
por: Latta-Lin, Daniel, et al.
Publicado: (2024) -
SGEMM-cube: Precision-Recovery FP32 GEMM Approximation on Ascend NPUs with FP16 Matrix Engines
por: Xue, Weicheng, et al.
Publicado: (2025) -
LAPIS: A Performance Portable, High Productivity Compiler Framework
por: Kelley, Brian, et al.
Publicado: (2025) -
W4A16 Mixed-Precision Matrix Multiplication on Decoupled Architecture: Kernel Design and Memory Bottleneck Analysis for Ascend NPUs
por: He, Yuanhong, et al.
Publicado: (2026) -
A-IO: Adaptive Inference Orchestration for Memory-Bound NPUs
por: Zhang, Chen, et al.
Publicado: (2026)