Context-Driven Performance Modeling for Causal Inference Operators on Neural Processing Units

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gupta, Neelesh, Jayanth, Rakshith, Parikh, Dhruv, Prasanna, Viktor
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912770281177088
author Gupta, Neelesh
Jayanth, Rakshith
Parikh, Dhruv
Prasanna, Viktor
author_facet Gupta, Neelesh
Jayanth, Rakshith
Parikh, Dhruv
Prasanna, Viktor
contents The proliferation of large language models has driven demand for long-context inference on resource-constrained edge platforms. However, deploying these models on Neural Processing Units (NPUs) presents significant challenges due to architectural mismatch: the quadratic complexity of standard attention conflicts with NPU memory and compute patterns. This paper presents a comprehensive performance analysis of causal inference operators on a modern NPU, benchmarking quadratic attention against sub-quadratic alternatives including structured state-space models and causal convolutions. Our analysis reveals a spectrum of critical bottlenecks: quadratic attention becomes severely memory-bound with catastrophic cache inefficiency, while sub-quadratic variants span from compute-bound on programmable vector cores to memory-bound by data movement. These findings provide essential insights for co-designing hardware-aware models and optimization strategies to enable efficient long-context inference on edge platforms.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25155
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Context-Driven Performance Modeling for Causal Inference Operators on Neural Processing Units
Gupta, Neelesh
Jayanth, Rakshith
Parikh, Dhruv
Prasanna, Viktor
Distributed, Parallel, and Cluster Computing
Machine Learning
The proliferation of large language models has driven demand for long-context inference on resource-constrained edge platforms. However, deploying these models on Neural Processing Units (NPUs) presents significant challenges due to architectural mismatch: the quadratic complexity of standard attention conflicts with NPU memory and compute patterns. This paper presents a comprehensive performance analysis of causal inference operators on a modern NPU, benchmarking quadratic attention against sub-quadratic alternatives including structured state-space models and causal convolutions. Our analysis reveals a spectrum of critical bottlenecks: quadratic attention becomes severely memory-bound with catastrophic cache inefficiency, while sub-quadratic variants span from compute-bound on programmable vector cores to memory-bound by data movement. These findings provide essential insights for co-designing hardware-aware models and optimization strategies to enable efficient long-context inference on edge platforms.
title Context-Driven Performance Modeling for Causal Inference Operators on Neural Processing Units
topic Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2509.25155