Flex Attention: A Programming Model for Generating Optimized Attention Kernels
Fuente:
arXiv
Saved in:
| Main Authors: | Dong, Juechu, Feng, Boyuan, Guessous, Driss, Liang, Yanbo, He, Horace |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FLuRKA: Fast and accurate unified Low-Rank & Kernel Attention
by: Gupta, Ahan, et al.
Published: (2023)
by: Gupta, Ahan, et al.
Published: (2023)
LOOPerSet: A Large-Scale Dataset for Data-Driven Polyhedral Compiler Optimization
by: Merouani, Massinissa, et al.
Published: (2025)
by: Merouani, Massinissa, et al.
Published: (2025)
The Next 700 ML-Enabled Compiler Optimizations
by: VenkataKeerthy, S., et al.
Published: (2023)
by: VenkataKeerthy, S., et al.
Published: (2023)
AttentionEngine: A Versatile Framework for Efficient Attention Mechanisms on Diverse Hardware Platforms
by: Chen, Feiyang, et al.
Published: (2025)
by: Chen, Feiyang, et al.
Published: (2025)
ECCO: Evidence-Driven Causal Reasoning for Compiler Optimization
by: Pan, Haolin, et al.
Published: (2026)
by: Pan, Haolin, et al.
Published: (2026)
Block Sparse Flash Attention
by: Ohayon, Daniel, et al.
Published: (2025)
by: Ohayon, Daniel, et al.
Published: (2025)
DaCe AD: Unifying High-Performance Automatic Differentiation for Machine Learning and Scientific Computing
by: Boudaoud, Afif, et al.
Published: (2025)
by: Boudaoud, Afif, et al.
Published: (2025)
Integration of a systolic array based hardware accelerator into a DNN operator auto-tuning framework
by: Peccia, F. N., et al.
Published: (2022)
by: Peccia, F. N., et al.
Published: (2022)
Insum: Sparse GPU Kernels Simplified and Optimized with Indirect Einsums
by: Won, Jaeyeon, et al.
Published: (2025)
by: Won, Jaeyeon, et al.
Published: (2025)
Who Wins the Race? (R Vs Python) - An Exploratory Study on Energy Consumption of Machine Learning Algorithms
by: Chattaraj, Rajrupa, et al.
Published: (2025)
by: Chattaraj, Rajrupa, et al.
Published: (2025)
Library Liberation: Competitive Performance Matmul Through Compiler-composed Nanokernels
by: Thangamani, Arun, et al.
Published: (2025)
by: Thangamani, Arun, et al.
Published: (2025)
Agentic Auto-Scheduling: An Experimental Study of LLM-Guided Loop Optimization
by: Merouani, Massinissa, et al.
Published: (2025)
by: Merouani, Massinissa, et al.
Published: (2025)
Ragged Paged Attention: A High-Performance and Flexible LLM Inference Kernel for TPU
by: Jiang, Jevin, et al.
Published: (2026)
by: Jiang, Jevin, et al.
Published: (2026)
Systematic Evaluation of Optimization Techniques for Long-Context Language Models
by: Ahmed, Ammar, et al.
Published: (2025)
by: Ahmed, Ammar, et al.
Published: (2025)
KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels
by: Wang, Han, et al.
Published: (2026)
by: Wang, Han, et al.
Published: (2026)
AutoKernel: Autonomous GPU Kernel Optimization via Iterative Agent-Driven Search
by: Jaber, Jaber, et al.
Published: (2026)
by: Jaber, Jaber, et al.
Published: (2026)
It's Not Easy Being Green: On the Energy Efficiency of Programming Languages
by: van Kempen, Nicolas, et al.
Published: (2024)
by: van Kempen, Nicolas, et al.
Published: (2024)
Automating Energy-Efficient GPU Kernel Generation: A Fast Search-Based Compilation Approach
by: Zhang, Yijia, et al.
Published: (2024)
by: Zhang, Yijia, et al.
Published: (2024)
Worst-Case Convergence Time of ML Algorithms via Extreme Value Theory
by: Tizpaz-Niari, Saeid, et al.
Published: (2024)
by: Tizpaz-Niari, Saeid, et al.
Published: (2024)
Stencil-Lifting: Hierarchical Recursive Lifting System for Extracting Summary of Stencil Kernel in Legacy Codes
by: Li, Mingyi, et al.
Published: (2025)
by: Li, Mingyi, et al.
Published: (2025)
Rule-Based Graph Programs Matching the Time Complexity of Imperative Algorithms
by: Alaoui, Ziad Ismaili, et al.
Published: (2025)
by: Alaoui, Ziad Ismaili, et al.
Published: (2025)
Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants
by: You, Bozhi, et al.
Published: (2025)
by: You, Bozhi, et al.
Published: (2025)
Optimizing Layout of Recursive Datatypes with Marmoset
by: Singhal, Vidush, et al.
Published: (2024)
by: Singhal, Vidush, et al.
Published: (2024)
FlashOmni: A Unified Sparse Attention Engine for Diffusion Transformers
by: Qiao, Liang, et al.
Published: (2025)
by: Qiao, Liang, et al.
Published: (2025)
Testing the Unknown: A Framework for OpenMP Testing via Random Program Generation
by: Laguna, Ignacio, et al.
Published: (2024)
by: Laguna, Ignacio, et al.
Published: (2024)
Evaluating Compiler Optimization Impacts on zkVM Performance
by: Gassmann, Thomas, et al.
Published: (2025)
by: Gassmann, Thomas, et al.
Published: (2025)
Priority Sampling of Large Language Models for Compilers
by: Grubisic, Dejan, et al.
Published: (2024)
by: Grubisic, Dejan, et al.
Published: (2024)
DF-GNN: Dynamic Fusion Framework for Attention Graph Neural Networks on GPUs
by: Liu, Jiahui, et al.
Published: (2024)
by: Liu, Jiahui, et al.
Published: (2024)
OMPILOT: Harnessing Transformer Models for Auto Parallelization to Shared Memory Computing Paradigms
by: Bhattacharjee, Arijit, et al.
Published: (2025)
by: Bhattacharjee, Arijit, et al.
Published: (2025)
CodeRosetta: Pushing the Boundaries of Unsupervised Code Translation for Parallel Programming
by: TehraniJamsaz, Ali, et al.
Published: (2024)
by: TehraniJamsaz, Ali, et al.
Published: (2024)
HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing
by: Liu, Minghui, et al.
Published: (2024)
by: Liu, Minghui, et al.
Published: (2024)
Model Compression and Efficient Inference for Large Language Models: A Survey
by: Wang, Wenxiao, et al.
Published: (2024)
by: Wang, Wenxiao, et al.
Published: (2024)
AutoLALA: Automatic Loop Algebraic Locality Analysis for AI and HPC Kernels
by: Zhu, Yifan, et al.
Published: (2026)
by: Zhu, Yifan, et al.
Published: (2026)
Enhancing Inference Efficiency of Large Language Models: Investigating Optimization Strategies and Architectural Innovations
by: Tyukin, Georgy
Published: (2024)
by: Tyukin, Georgy
Published: (2024)
LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
by: Yang, Shang, et al.
Published: (2025)
by: Yang, Shang, et al.
Published: (2025)
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
by: Liu, Zirui, et al.
Published: (2024)
by: Liu, Zirui, et al.
Published: (2024)
AutoSAGE: Input-Aware CUDA Scheduling for Sparse GNN Aggregation (SpMM/SDDMM) and CSR Attention
by: Stankovic, Aleksandar
Published: (2025)
by: Stankovic, Aleksandar
Published: (2025)
Efficient Graph Knowledge Distillation from GNNs to Kolmogorov--Arnold Networks via Self-Attention Dynamic Sampling
by: Cui, Can, et al.
Published: (2025)
by: Cui, Can, et al.
Published: (2025)
Data Efficacy for Language Model Training
by: Dai, Yalun, et al.
Published: (2025)
by: Dai, Yalun, et al.
Published: (2025)
ModeSwitch-LLM: A Lightweight Phase-Aware Controller for Cross-Mode LLM Inference on a Single GPU
by: Sunesh, Aman, et al.
Published: (2026)
by: Sunesh, Aman, et al.
Published: (2026)
Similar Items
-
FLuRKA: Fast and accurate unified Low-Rank & Kernel Attention
by: Gupta, Ahan, et al.
Published: (2023) -
LOOPerSet: A Large-Scale Dataset for Data-Driven Polyhedral Compiler Optimization
by: Merouani, Massinissa, et al.
Published: (2025) -
The Next 700 ML-Enabled Compiler Optimizations
by: VenkataKeerthy, S., et al.
Published: (2023) -
AttentionEngine: A Versatile Framework for Efficient Attention Mechanisms on Diverse Hardware Platforms
by: Chen, Feiyang, et al.
Published: (2025) -
ECCO: Evidence-Driven Causal Reasoning for Compiler Optimization
by: Pan, Haolin, et al.
Published: (2026)