Hexagon-MLIR: An AI Compilation Stack For Qualcomm's Neural Processing Units (NPUs)

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Absar, Mohammed Javed, Baskaran, Muthu, Sharma, Abhikrant, Bhandari, Abhilash, Aggarwal, Ankit, Rangasamy, Arun, Das, Dibyendu, Hosseini, Fateme, Slama, Franck, Brumar, Iulian, Verma, Jyotsna, Bindumadhavan, Krishnaprasad, Kothari, Mitesh, Gupta, Mohit, Kolachana, Ravishankar, Lethin, Richard, Narang, Samarth, Ladwa, Sanjay Motilal, Jain, Shalini, Dalvi, Snigdha Suresh, Rahman, Tasmia, Komatireddy, Venkat Rasagna Reddy, Pandya, Vivek Vasudevbhai, Shi, Xiyue, Zipper, Zachary
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917288545878016
author Absar, Mohammed Javed
Baskaran, Muthu
Sharma, Abhikrant
Bhandari, Abhilash
Aggarwal, Ankit
Rangasamy, Arun
Das, Dibyendu
Hosseini, Fateme
Slama, Franck
Brumar, Iulian
Verma, Jyotsna
Bindumadhavan, Krishnaprasad
Kothari, Mitesh
Gupta, Mohit
Kolachana, Ravishankar
Lethin, Richard
Narang, Samarth
Ladwa, Sanjay Motilal
Jain, Shalini
Dalvi, Snigdha Suresh
Rahman, Tasmia
Komatireddy, Venkat Rasagna Reddy
Pandya, Vivek Vasudevbhai
Shi, Xiyue
Zipper, Zachary
author_facet Absar, Mohammed Javed
Baskaran, Muthu
Sharma, Abhikrant
Bhandari, Abhilash
Aggarwal, Ankit
Rangasamy, Arun
Das, Dibyendu
Hosseini, Fateme
Slama, Franck
Brumar, Iulian
Verma, Jyotsna
Bindumadhavan, Krishnaprasad
Kothari, Mitesh
Gupta, Mohit
Kolachana, Ravishankar
Lethin, Richard
Narang, Samarth
Ladwa, Sanjay Motilal
Jain, Shalini
Dalvi, Snigdha Suresh
Rahman, Tasmia
Komatireddy, Venkat Rasagna Reddy
Pandya, Vivek Vasudevbhai
Shi, Xiyue
Zipper, Zachary
contents In this paper, we present Hexagon-MLIR,an open-source compilation stack that targets Qualcomm Hexagon Neural Processing Unit (NPU) and provides unified support for lowering Triton kernels and PyTorch models . Built using the MLIR framework, our compiler applies a structured sequence of passes to exploit NPU architectural features to accelerate AI workloads. It enables faster deployment of new Triton kernels (hand-written or subgraphs from PyTorch 2.0), for our target by providing automated compilation from kernel to binary. By ingesting Triton kernels, we generate mega-kernels that maximize data locality in the NPU's Tightly Coupled Memory (TCM), reducing the bandwidth bottlenecks inherent in library-based approaches. This initiative complements our commercial toolchains by providing developers with an open-source MLIR-based compilation stack that gives them a path to advance AI compilation capabilities through a more flexible approach. Hexagon-MLIR is a work-in-progress, and we are continuing to add many more optimizations and capabilities in this effort.
format Preprint
id arxiv_https___arxiv_org_abs_2602_19762
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Hexagon-MLIR: An AI Compilation Stack For Qualcomm's Neural Processing Units (NPUs)
Absar, Mohammed Javed
Baskaran, Muthu
Sharma, Abhikrant
Bhandari, Abhilash
Aggarwal, Ankit
Rangasamy, Arun
Das, Dibyendu
Hosseini, Fateme
Slama, Franck
Brumar, Iulian
Verma, Jyotsna
Bindumadhavan, Krishnaprasad
Kothari, Mitesh
Gupta, Mohit
Kolachana, Ravishankar
Lethin, Richard
Narang, Samarth
Ladwa, Sanjay Motilal
Jain, Shalini
Dalvi, Snigdha Suresh
Rahman, Tasmia
Komatireddy, Venkat Rasagna Reddy
Pandya, Vivek Vasudevbhai
Shi, Xiyue
Zipper, Zachary
Programming Languages
Artificial Intelligence
In this paper, we present Hexagon-MLIR,an open-source compilation stack that targets Qualcomm Hexagon Neural Processing Unit (NPU) and provides unified support for lowering Triton kernels and PyTorch models . Built using the MLIR framework, our compiler applies a structured sequence of passes to exploit NPU architectural features to accelerate AI workloads. It enables faster deployment of new Triton kernels (hand-written or subgraphs from PyTorch 2.0), for our target by providing automated compilation from kernel to binary. By ingesting Triton kernels, we generate mega-kernels that maximize data locality in the NPU's Tightly Coupled Memory (TCM), reducing the bandwidth bottlenecks inherent in library-based approaches. This initiative complements our commercial toolchains by providing developers with an open-source MLIR-based compilation stack that gives them a path to advance AI compilation capabilities through a more flexible approach. Hexagon-MLIR is a work-in-progress, and we are continuing to add many more optimizations and capabilities in this effort.
title Hexagon-MLIR: An AI Compilation Stack For Qualcomm's Neural Processing Units (NPUs)
topic Programming Languages
Artificial Intelligence
url https://arxiv.org/abs/2602.19762