A Performance Evaluation of a Quantized Large Language Model on Various Smartphones

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Çöplü, Tolga, Loedi, Marc, Bendiken, Arto, Makohin, Mykhailo, Bouw, Joshua J., Cobb, Stephen
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911768669847552
author Çöplü, Tolga
Loedi, Marc
Bendiken, Arto
Makohin, Mykhailo
Bouw, Joshua J.
Cobb, Stephen
author_facet Çöplü, Tolga
Loedi, Marc
Bendiken, Arto
Makohin, Mykhailo
Bouw, Joshua J.
Cobb, Stephen
contents This paper explores the feasibility and performance of on-device large language model (LLM) inference on various Apple iPhone models. Amidst the rapid evolution of generative AI, on-device LLMs offer solutions to privacy, security, and connectivity challenges inherent in cloud-based models. Leveraging existing literature on running multi-billion parameter LLMs on resource-limited devices, our study examines the thermal effects and interaction speeds of a high-performing LLM across different smartphone generations. We present real-world performance results, providing insights into on-device inference capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2312_12472
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle A Performance Evaluation of a Quantized Large Language Model on Various Smartphones
Çöplü, Tolga
Loedi, Marc
Bendiken, Arto
Makohin, Mykhailo
Bouw, Joshua J.
Cobb, Stephen
Machine Learning
Artificial Intelligence
Performance
I.2.7
This paper explores the feasibility and performance of on-device large language model (LLM) inference on various Apple iPhone models. Amidst the rapid evolution of generative AI, on-device LLMs offer solutions to privacy, security, and connectivity challenges inherent in cloud-based models. Leveraging existing literature on running multi-billion parameter LLMs on resource-limited devices, our study examines the thermal effects and interaction speeds of a high-performing LLM across different smartphone generations. We present real-world performance results, providing insights into on-device inference capabilities.
title A Performance Evaluation of a Quantized Large Language Model on Various Smartphones
topic Machine Learning
Artificial Intelligence
Performance
I.2.7
url https://arxiv.org/abs/2312.12472