NITRO: LLM Inference on Intel Laptop NPUs
Fuente:
arXiv
Saved in:
| Main Authors: | Fei, Anthony, Abdelfattah, Mohamed S. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unleashing Low-Bit Inference on Ascend NPUs: A Comprehensive Evaluation of HiFloat Formats
by: Zhao, Pengxiang, et al.
Published: (2026)
by: Zhao, Pengxiang, et al.
Published: (2026)
Pangu Ultra: Pushing the Limits of Dense Large Language Models on Ascend NPUs
by: Yin, Yichun, et al.
Published: (2025)
by: Yin, Yichun, et al.
Published: (2025)
Fast On-device LLM Inference with NPUs
by: Xu, Daliang, et al.
Published: (2024)
by: Xu, Daliang, et al.
Published: (2024)
HiFloat4 Format for Language Model Pre-training on Ascend NPUs
by: Taghian, Mehran, et al.
Published: (2026)
by: Taghian, Mehran, et al.
Published: (2026)
ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models
by: Akhauri, Yash, et al.
Published: (2024)
by: Akhauri, Yash, et al.
Published: (2024)
Star Attention: Efficient LLM Inference over Long Sequences
by: Acharya, Shantanu, et al.
Published: (2024)
by: Acharya, Shantanu, et al.
Published: (2024)
MemeIntel: Explainable Detection of Propagandistic and Hateful Memes
by: Kmainasi, Mohamed Bayan, et al.
Published: (2025)
by: Kmainasi, Mohamed Bayan, et al.
Published: (2025)
Propensity Inference: Environmental Contributors to LLM Behaviour
by: Järviniemi, Olli, et al.
Published: (2026)
by: Järviniemi, Olli, et al.
Published: (2026)
LLM Inference Unveiled: Survey and Roofline Model Insights
by: Yuan, Zhihang, et al.
Published: (2024)
by: Yuan, Zhihang, et al.
Published: (2024)
Scaling LLM Inference with Optimized Sample Compute Allocation
by: Zhang, Kexun, et al.
Published: (2024)
by: Zhang, Kexun, et al.
Published: (2024)
LLM-based Translation Inference with Iterative Bilingual Understanding
by: Chen, Andong, et al.
Published: (2024)
by: Chen, Andong, et al.
Published: (2024)
Hexagon-MLIR: An AI Compilation Stack For Qualcomm's Neural Processing Units (NPUs)
by: Absar, Mohammed Javed, et al.
Published: (2026)
by: Absar, Mohammed Javed, et al.
Published: (2026)
EvoP: Robust LLM Inference via Evolutionary Pruning
by: Wu, Shangyu, et al.
Published: (2025)
by: Wu, Shangyu, et al.
Published: (2025)
A-IO: Adaptive Inference Orchestration for Memory-Bound NPUs
by: Zhang, Chen, et al.
Published: (2026)
by: Zhang, Chen, et al.
Published: (2026)
XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference
by: Monteiro, João, et al.
Published: (2024)
by: Monteiro, João, et al.
Published: (2024)
Amphista: Bi-directional Multi-head Decoding for Accelerating LLM Inference
by: Li, Zeping, et al.
Published: (2024)
by: Li, Zeping, et al.
Published: (2024)
Reformulating KV Cache Eviction Problem for Long-Context LLM Inference
by: Mai, Tho, et al.
Published: (2026)
by: Mai, Tho, et al.
Published: (2026)
FLRC: Fine-grained Low-Rank Compressor for Efficient LLM Inference
by: Lu, Yu-Chen, et al.
Published: (2025)
by: Lu, Yu-Chen, et al.
Published: (2025)
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference
by: Ouyang, Haojie, et al.
Published: (2025)
by: Ouyang, Haojie, et al.
Published: (2025)
From Natural Language to SQL: Review of LLM-based Text-to-SQL Systems
by: Mohammadjafari, Ali, et al.
Published: (2024)
by: Mohammadjafari, Ali, et al.
Published: (2024)
Is Multilingual LLM Watermarking Truly Multilingual? Scaling Robustness to 100+ Languages via Back-Translation
by: Mohamed, Asim, et al.
Published: (2025)
by: Mohamed, Asim, et al.
Published: (2025)
Can We Trust LLM Detectors?
by: Sandhan, Jivnesh, et al.
Published: (2026)
by: Sandhan, Jivnesh, et al.
Published: (2026)
MindVL: Towards Efficient and Effective Training of Multimodal Large Language Models on Ascend NPUs
by: Chen, Feilong, et al.
Published: (2025)
by: Chen, Feilong, et al.
Published: (2025)
Regression Language Models for Code
by: Akhauri, Yash, et al.
Published: (2025)
by: Akhauri, Yash, et al.
Published: (2025)
AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM Inference
by: He, Zhuomin, et al.
Published: (2025)
by: He, Zhuomin, et al.
Published: (2025)
Performance Characterization of Expert Router for Scalable LLM Inference
by: Pichlmeier, Josef, et al.
Published: (2024)
by: Pichlmeier, Josef, et al.
Published: (2024)
DeFT: Decoding with Flash Tree-attention for Efficient Tree-structured LLM Inference
by: Yao, Jinwei, et al.
Published: (2024)
by: Yao, Jinwei, et al.
Published: (2024)
A Survey of Frontiers in LLM Reasoning: Inference Scaling, Learning to Reason, and Agentic Systems
by: Ke, Zixuan, et al.
Published: (2025)
by: Ke, Zixuan, et al.
Published: (2025)
DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference
by: Dehghanighobadi, Zahra, et al.
Published: (2026)
by: Dehghanighobadi, Zahra, et al.
Published: (2026)
Round Attention: A Novel Round-Level Attention Mechanism to Accelerate LLM Inference
by: Tang, Yaohua, et al.
Published: (2025)
by: Tang, Yaohua, et al.
Published: (2025)
OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference
by: Gu, Yuzhe, et al.
Published: (2025)
by: Gu, Yuzhe, et al.
Published: (2025)
A Simple Ensemble Strategy for LLM Inference: Towards More Stable Text Classification
by: Niimi, Junichiro
Published: (2025)
by: Niimi, Junichiro
Published: (2025)
Epistemic Blinding: An Inference-Time Protocol for Auditing Prior Contamination in LLM-Assisted Analysis
by: Cuccarese, Michael
Published: (2026)
by: Cuccarese, Michael
Published: (2026)
LLM as a Broken Telephone: Iterative Generation Distorts Information
by: Mohamed, Amr, et al.
Published: (2025)
by: Mohamed, Amr, et al.
Published: (2025)
Exploring Spatial Representations in the Historical Lake District Texts with LLM-based Relation Extraction
by: Haris, Erum, et al.
Published: (2024)
by: Haris, Erum, et al.
Published: (2024)
Bridging Writing Manner Gap in Visual Instruction Tuning by Creating LLM-aligned Instructions
by: Jing, Dong, et al.
Published: (2025)
by: Jing, Dong, et al.
Published: (2025)
STRUX: An LLM for Decision-Making with Structured Explanations
by: Lu, Yiming, et al.
Published: (2024)
by: Lu, Yiming, et al.
Published: (2024)
TokenButler: Token Importance is Predictable
by: Akhauri, Yash, et al.
Published: (2025)
by: Akhauri, Yash, et al.
Published: (2025)
Why Do LLM-based Web Agents Fail? A Hierarchical Planning Perspective
by: Aghzal, Mohamed, et al.
Published: (2026)
by: Aghzal, Mohamed, et al.
Published: (2026)
Active Inference for Self-Organizing Multi-LLM Systems: A Bayesian Thermodynamic Approach to Adaptation
by: Prakki, Rithvik
Published: (2024)
by: Prakki, Rithvik
Published: (2024)
Similar Items
-
Unleashing Low-Bit Inference on Ascend NPUs: A Comprehensive Evaluation of HiFloat Formats
by: Zhao, Pengxiang, et al.
Published: (2026) -
Pangu Ultra: Pushing the Limits of Dense Large Language Models on Ascend NPUs
by: Yin, Yichun, et al.
Published: (2025) -
Fast On-device LLM Inference with NPUs
by: Xu, Daliang, et al.
Published: (2024) -
HiFloat4 Format for Language Model Pre-training on Ascend NPUs
by: Taghian, Mehran, et al.
Published: (2026) -
ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models
by: Akhauri, Yash, et al.
Published: (2024)