Generative AI on the Edge: Architecture and Performance Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nezami, Zeinab, Hafeez, Maryam, Djemame, Karim, Zaidi, Syed Ali Raza
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929606158712832
author Nezami, Zeinab
Hafeez, Maryam
Djemame, Karim
Zaidi, Syed Ali Raza
author_facet Nezami, Zeinab
Hafeez, Maryam
Djemame, Karim
Zaidi, Syed Ali Raza
contents 6G's AI native vision of embedding advance intelligence in the network while bringing it closer to the user requires a systematic evaluation of Generative AI (GenAI) models on edge devices. Rapidly emerging solutions based on Open RAN (ORAN) and Network-in-a-Box strongly advocate the use of low-cost, off-the-shelf components for simpler and efficient deployment, e.g., in provisioning rural connectivity. In this context, conceptual architecture, hardware testbeds and precise performance quantification of Large Language Models (LLMs) on off-the-shelf edge devices remains largely unexplored. This research investigates computationally demanding LLM inference on a single commodity Raspberry Pi serving as an edge testbed for ORAN. We investigate various LLMs, including small, medium and large models, on a Raspberry Pi 5 Cluster using a lightweight Kubernetes distribution (K3s) with modular prompting implementation. We study its feasibility and limitations by analyzing throughput, latency, accuracy and efficiency. Our findings indicate that CPU-only deployment of lightweight models, such as Yi, Phi, and Llama3, can effectively support edge applications, achieving a generation throughput of 5 to 12 tokens per second with less than 50\% CPU and RAM usage. We conclude that GenAI on the edge offers localized inference in remote or bandwidth-constrained environments in 6G networks without reliance on cloud infrastructure.
format Preprint
id arxiv_https___arxiv_org_abs_2411_17712
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Generative AI on the Edge: Architecture and Performance Evaluation
Nezami, Zeinab
Hafeez, Maryam
Djemame, Karim
Zaidi, Syed Ali Raza
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Networking and Internet Architecture
Performance
6G's AI native vision of embedding advance intelligence in the network while bringing it closer to the user requires a systematic evaluation of Generative AI (GenAI) models on edge devices. Rapidly emerging solutions based on Open RAN (ORAN) and Network-in-a-Box strongly advocate the use of low-cost, off-the-shelf components for simpler and efficient deployment, e.g., in provisioning rural connectivity. In this context, conceptual architecture, hardware testbeds and precise performance quantification of Large Language Models (LLMs) on off-the-shelf edge devices remains largely unexplored. This research investigates computationally demanding LLM inference on a single commodity Raspberry Pi serving as an edge testbed for ORAN. We investigate various LLMs, including small, medium and large models, on a Raspberry Pi 5 Cluster using a lightweight Kubernetes distribution (K3s) with modular prompting implementation. We study its feasibility and limitations by analyzing throughput, latency, accuracy and efficiency. Our findings indicate that CPU-only deployment of lightweight models, such as Yi, Phi, and Llama3, can effectively support edge applications, achieving a generation throughput of 5 to 12 tokens per second with less than 50\% CPU and RAM usage. We conclude that GenAI on the edge offers localized inference in remote or bandwidth-constrained environments in 6G networks without reliance on cloud infrastructure.
title Generative AI on the Edge: Architecture and Performance Evaluation
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Networking and Internet Architecture
Performance
url https://arxiv.org/abs/2411.17712