Neuromorphic Principles for Efficient Large Language Models on Intel Loihi 2

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Abreu, Steven, Shrestha, Sumit Bam, Zhu, Rui-Jie, Eshraghian, Jason
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916661752233984
author Abreu, Steven
Shrestha, Sumit Bam
Zhu, Rui-Jie
Eshraghian, Jason
author_facet Abreu, Steven
Shrestha, Sumit Bam
Zhu, Rui-Jie
Eshraghian, Jason
contents Large language models (LLMs) deliver impressive performance but require large amounts of energy. In this work, we present a MatMul-free LLM architecture adapted for Intel's neuromorphic processor, Loihi 2. Our approach leverages Loihi 2's support for low-precision, event-driven computation and stateful processing. Our hardware-aware quantized model on GPU demonstrates that a 370M parameter MatMul-free model can be quantized with no accuracy loss. Based on preliminary results, we report up to 3x higher throughput with 2x less energy, compared to transformer-based LLMs on an edge GPU, with significantly better scaling. Further hardware optimizations will increase throughput and decrease energy consumption. These results show the potential of neuromorphic hardware for efficient inference and pave the way for efficient reasoning models capable of generating complex, long-form text rapidly and cost-effectively.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18002
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Neuromorphic Principles for Efficient Large Language Models on Intel Loihi 2
Abreu, Steven
Shrestha, Sumit Bam
Zhu, Rui-Jie
Eshraghian, Jason
Neural and Evolutionary Computing
Artificial Intelligence
Hardware Architecture
Machine Learning
Large language models (LLMs) deliver impressive performance but require large amounts of energy. In this work, we present a MatMul-free LLM architecture adapted for Intel's neuromorphic processor, Loihi 2. Our approach leverages Loihi 2's support for low-precision, event-driven computation and stateful processing. Our hardware-aware quantized model on GPU demonstrates that a 370M parameter MatMul-free model can be quantized with no accuracy loss. Based on preliminary results, we report up to 3x higher throughput with 2x less energy, compared to transformer-based LLMs on an edge GPU, with significantly better scaling. Further hardware optimizations will increase throughput and decrease energy consumption. These results show the potential of neuromorphic hardware for efficient inference and pave the way for efficient reasoning models capable of generating complex, long-form text rapidly and cost-effectively.
title Neuromorphic Principles for Efficient Large Language Models on Intel Loihi 2
topic Neural and Evolutionary Computing
Artificial Intelligence
Hardware Architecture
Machine Learning
url https://arxiv.org/abs/2503.18002