PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xue, Zhenliang, Song, Yixin, Mi, Zeyu, Zheng, Xinrui, Xia, Yubin, Chen, Haibo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912152620630016
author Xue, Zhenliang
Song, Yixin
Mi, Zeyu
Zheng, Xinrui
Xia, Yubin
Chen, Haibo
author_facet Xue, Zhenliang
Song, Yixin
Mi, Zeyu
Zheng, Xinrui
Xia, Yubin
Chen, Haibo
contents Large language models (LLMs) on smartphones enable real-time AI assistance and privacy-preserving, offline operation. However, resource constraints of smartphones limit current deployments to small language models (SLMs), significantly compromising their capabilities. This paper introduces PowerInfer-2, a smartphone-based framework that enables fast inference for LLMs exceeding the memory capacity. The key insight is decomposing matrix operations into neuron clusters as the basic processing unit, which enables flexible scheduling and efficient I/O-computation pipelining. PowerInfer-2 leverages this neuron-cluster-based design in both computation and storage. For computation, neuron clusters with dense activations are processed on NPU, while sparse clusters use CPU. The storage engine provides a fine-grained pipeline mechanism that coordinates cluster-level computation and I/O operations, enhanced by a segmented neuron cache to reduce I/O activities. PowerInfer-2 achieves up to a 27.8x speed increase compared to state-of-the-art frameworks. PowerInfer-2 is the first system to serve a 47B LLM on a smartphone, achieving 11.68 tokens/s. Notably, these performance improvements preserve model quality with negligible accuracy degradation.
format Preprint
id arxiv_https___arxiv_org_abs_2406_06282
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Xue, Zhenliang
Song, Yixin
Mi, Zeyu
Zheng, Xinrui
Xia, Yubin
Chen, Haibo
Machine Learning
Large language models (LLMs) on smartphones enable real-time AI assistance and privacy-preserving, offline operation. However, resource constraints of smartphones limit current deployments to small language models (SLMs), significantly compromising their capabilities. This paper introduces PowerInfer-2, a smartphone-based framework that enables fast inference for LLMs exceeding the memory capacity. The key insight is decomposing matrix operations into neuron clusters as the basic processing unit, which enables flexible scheduling and efficient I/O-computation pipelining. PowerInfer-2 leverages this neuron-cluster-based design in both computation and storage. For computation, neuron clusters with dense activations are processed on NPU, while sparse clusters use CPU. The storage engine provides a fine-grained pipeline mechanism that coordinates cluster-level computation and I/O operations, enhanced by a segmented neuron cache to reduce I/O activities. PowerInfer-2 achieves up to a 27.8x speed increase compared to state-of-the-art frameworks. PowerInfer-2 is the first system to serve a 47B LLM on a smartphone, achieving 11.68 tokens/s. Notably, these performance improvements preserve model quality with negligible accuracy degradation.
title PowerInfer-2: Fast Large Language Model Inference on a Smartphone
topic Machine Learning
url https://arxiv.org/abs/2406.06282