Efficiency, Expressivity, and Extensibility in a Close-to-Metal NPU Programming Interface
Fuente:
arXiv
Saved in:
| Main Authors: | Hunhoff, Erika, Melber, Joseph, Denolf, Kristof, Bisca, Andra, Bayliss, Samuel, Neuendorffer, Stephen, Fifield, Jeff, Lo, Jack, Vasireddy, Pranathi, James-Roxby, Phil, Keller, Eric |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Loop Nests to Silicon: Mapping AI Workloads onto AMD NPUs with MLIR-AIR
by: Wang, Erwei, et al.
Published: (2025)
by: Wang, Erwei, et al.
Published: (2025)
Striking the Balance: GEMM Performance Optimization Across Generations of Ryzen AI NPUs
by: Taka, Endri, et al.
Published: (2025)
by: Taka, Endri, et al.
Published: (2025)
AdaHOP: Fast and Accurate Low-Precision Training via Outlier-Pattern-Aware Rotation
by: Kim, Seonggon, et al.
Published: (2026)
by: Kim, Seonggon, et al.
Published: (2026)
Practical Formal Verification for MLIR Programs
by: Tucker, Emily, et al.
Published: (2026)
by: Tucker, Emily, et al.
Published: (2026)
Error Diffusion: Post Training Quantization with Block-Scaled Number Formats for Neural Networks
by: Khodamoradi, Alireza, et al.
Published: (2024)
by: Khodamoradi, Alireza, et al.
Published: (2024)
Colonialism, Genocide and Reparations: The German‐Namibian Case
by: Henning Melber
Published: (2024)
by: Henning Melber
Published: (2024)
Can Asymmetric Tile Buffering Be Beneficial?
by: Wang, Chengyue, et al.
Published: (2025)
by: Wang, Chengyue, et al.
Published: (2025)
JUROS SOBRE CAPITAL PRÓPRIO: UMA ANÁLISE SOBRE O IMPACTO TRIBUTÁRIO PARA QUEM PAGA E PARA QUEM RECEBE
by: Maria Heloisa Bisca
Published: (2012)
by: Maria Heloisa Bisca
Published: (2012)
LoRaQ: Optimized Low Rank Approximation for 4-bit Quantization
by: Bouquet, Yann, et al.
Published: (2026)
by: Bouquet, Yann, et al.
Published: (2026)
Contra la corriente
by: Fifield, A
Published: (2008)
by: Fifield, A
Published: (2008)
Never Too Old for Stories (Tales of a High School Information Specialist).
by: Fifield, Carol
Published: (1996)
by: Fifield, Carol
Published: (1996)
Evaluating the Energy Efficiency of NPU-Accelerated Machine Learning Inference on Embedded Microcontrollers
by: Fanariotis, Anastasios, et al.
Published: (2025)
by: Fanariotis, Anastasios, et al.
Published: (2025)
ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference
by: Yin, Wangsong, et al.
Published: (2025)
by: Yin, Wangsong, et al.
Published: (2025)
Evaluation of a Career Development and Assessment Center Program for Professional Librarians.
by: Melber, Barbara D., et al.
Published: (1985)
by: Melber, Barbara D., et al.
Published: (1985)
Express4D: Expressive, Friendly, and Extensible 4D Facial Motion Generation Benchmark
by: Aloni, Yaron, et al.
Published: (2025)
by: Aloni, Yaron, et al.
Published: (2025)
EdgeInfinite-Instruct: Bridging SFT-Based Optimization and NPU-Level Efficiency for Edge Devices
by: Chen, Jiyu, et al.
Published: (2025)
by: Chen, Jiyu, et al.
Published: (2025)
Evaluating ChatGPT's Performance in Classifying Pneumonia from Chest X-Ray Images
by: Prahallad, Pragna, et al.
Published: (2025)
by: Prahallad, Pragna, et al.
Published: (2025)
LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load
by: Tummalapalli, Pranay, et al.
Published: (2026)
by: Tummalapalli, Pranay, et al.
Published: (2026)
Fleet: Hierarchical Task-based Abstraction for Megakernels on Multi-Die GPUs
by: Chowdhary, Sangeeta, et al.
Published: (2026)
by: Chowdhary, Sangeeta, et al.
Published: (2026)
RUBICON: A Framework for Designing Efficient Deep Learning-Based Genomic Basecallers
by: Singh, Gagandeep, et al.
Published: (2022)
by: Singh, Gagandeep, et al.
Published: (2022)
Using an interactive simulation tool of national energy and climate policy planning to explore environmental policy options from the perspectives of different interest groups
by: Andra Blumberga
Published: (2024)
by: Andra Blumberga
Published: (2024)
Subject case alternation in negated existential, locative, and possessive clauses in Latvian
by: Andra Kalnača
Published: (2018)
by: Andra Kalnača
Published: (2018)
Open WebUI: An Open, Extensible, and Usable Interface for AI Interaction
by: Baek, Jaeryang, et al.
Published: (2025)
by: Baek, Jaeryang, et al.
Published: (2025)
NPU Design for Diffusion Language Model Inference
by: Lou, Binglei, et al.
Published: (2026)
by: Lou, Binglei, et al.
Published: (2026)
IRFuzzer: Specialized Fuzzing for LLVM Backend Code Generation
by: Rong, Yuyang, et al.
Published: (2024)
by: Rong, Yuyang, et al.
Published: (2024)
WindVE: Collaborative CPU-NPU Vector Embedding
by: Huang, Jinqi, et al.
Published: (2025)
by: Huang, Jinqi, et al.
Published: (2025)
Implementation and Optimization of HQC Decoding on NPU-Integrated Devices
by: Chau, Vu Minh, et al.
Published: (2026)
by: Chau, Vu Minh, et al.
Published: (2026)
NPU-NTU System for Voice Privacy 2024 Challenge
by: Yao, Jixun, et al.
Published: (2024)
by: Yao, Jixun, et al.
Published: (2024)
Towards Efficient Multi-Scale Deformable Attention on NPU
by: Huang, Chenghuan, et al.
Published: (2025)
by: Huang, Chenghuan, et al.
Published: (2025)
NTU-NPU System for Voice Privacy 2024 Challenge
by: Kuzmin, Nikita, et al.
Published: (2024)
by: Kuzmin, Nikita, et al.
Published: (2024)
Zapotec weavers of Teotitl n / Andra Fischgrund Stanton ; field photographs by Jay Phillips
by: Stanton, Andra Fischgrund
by: Stanton, Andra Fischgrund
Zermelian Extensibility
by: Andrew Bacon
Published: (2026)
by: Andrew Bacon
Published: (2026)
Expressivity-Efficiency Tradeoffs for Hybrid Sequence Models
by: Cooper, John, et al.
Published: (2026)
by: Cooper, John, et al.
Published: (2026)
Multi-View Oriented GPLVM: Expressiveness and Efficiency
by: Yang, Zi, et al.
Published: (2025)
by: Yang, Zi, et al.
Published: (2025)
Expressive Symbolic Regression for Interpretable Models of Discrete-Time Dynamical Systems
by: Iyer, Adarsh, et al.
Published: (2024)
by: Iyer, Adarsh, et al.
Published: (2024)
Implementing Keyword Spotting on the MCUX947 Microcontroller with Integrated NPU
by: Jakuš, Petar, et al.
Published: (2025)
by: Jakuš, Petar, et al.
Published: (2025)
EONSim: An NPU Simulator for On-Chip Memory and Embedding Vector Operations
by: Choi, Sangun, et al.
Published: (2025)
by: Choi, Sangun, et al.
Published: (2025)
Accelerating OpenPangu Inference on NPU via Speculative Decoding
by: Dai, Yuntao, et al.
Published: (2026)
by: Dai, Yuntao, et al.
Published: (2026)
NPUEval: Optimizing NPU Kernels with LLMs and Open Source Compilers
by: Kalade, Sarunas, et al.
Published: (2025)
by: Kalade, Sarunas, et al.
Published: (2025)
Scaling LLM Test-Time Compute with Mobile NPU on Smartphones
by: Hao, Zixu, et al.
Published: (2025)
by: Hao, Zixu, et al.
Published: (2025)
Similar Items
-
From Loop Nests to Silicon: Mapping AI Workloads onto AMD NPUs with MLIR-AIR
by: Wang, Erwei, et al.
Published: (2025) -
Striking the Balance: GEMM Performance Optimization Across Generations of Ryzen AI NPUs
by: Taka, Endri, et al.
Published: (2025) -
AdaHOP: Fast and Accurate Low-Precision Training via Outlier-Pattern-Aware Rotation
by: Kim, Seonggon, et al.
Published: (2026) -
Practical Formal Verification for MLIR Programs
by: Tucker, Emily, et al.
Published: (2026) -
Error Diffusion: Post Training Quantization with Block-Scaled Number Formats for Neural Networks
by: Khodamoradi, Alireza, et al.
Published: (2024)