Characterizing and Understanding Energy Footprint and Efficiency of Small Language Model on Edges

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Islam, Md Romyull, Deng, Bobin, Dhar, Nobel, Nguyen, Tu N., He, Selena, Shi, Yong, Suo, Kun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914158857945088
author Islam, Md Romyull
Deng, Bobin
Dhar, Nobel
Nguyen, Tu N.
He, Selena
Shi, Yong
Suo, Kun
author_facet Islam, Md Romyull
Deng, Bobin
Dhar, Nobel
Nguyen, Tu N.
He, Selena
Shi, Yong
Suo, Kun
contents Cloud-based large language models (LLMs) and their variants have significantly influenced real-world applications. Deploying smaller models (i.e., small language models (SLMs)) on edge devices offers additional advantages, such as reduced latency and independence from network connectivity. However, edge devices' limited computing resources and constrained energy budgets challenge efficient deployment. This study evaluates the power efficiency of five representative SLMs - Llama 3.2, Phi-3 Mini, TinyLlama, and Gemma 2 on Raspberry Pi 5, Jetson Nano, and Jetson Orin Nano (CPU and GPU configurations). Results show that Jetson Orin Nano with GPU acceleration achieves the highest energy-to-performance ratio, significantly outperforming CPU-based setups. Llama 3.2 provides the best balance of accuracy and power efficiency, while TinyLlama is well-suited for low-power environments at the cost of reduced accuracy. In contrast, Phi-3 Mini consumes the most energy despite its high accuracy. In addition, GPU acceleration, memory bandwidth, and model architecture are key in optimizing inference energy efficiency. Our empirical analysis offers practical insights for AI, smart systems, and mobile ad-hoc platforms to leverage tradeoffs from accuracy, inference latency, and power efficiency in energy-constrained environments.
format Preprint
id arxiv_https___arxiv_org_abs_2511_11624
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Characterizing and Understanding Energy Footprint and Efficiency of Small Language Model on Edges
Islam, Md Romyull
Deng, Bobin
Dhar, Nobel
Nguyen, Tu N.
He, Selena
Shi, Yong
Suo, Kun
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Computation and Language
Machine Learning
Cloud-based large language models (LLMs) and their variants have significantly influenced real-world applications. Deploying smaller models (i.e., small language models (SLMs)) on edge devices offers additional advantages, such as reduced latency and independence from network connectivity. However, edge devices' limited computing resources and constrained energy budgets challenge efficient deployment. This study evaluates the power efficiency of five representative SLMs - Llama 3.2, Phi-3 Mini, TinyLlama, and Gemma 2 on Raspberry Pi 5, Jetson Nano, and Jetson Orin Nano (CPU and GPU configurations). Results show that Jetson Orin Nano with GPU acceleration achieves the highest energy-to-performance ratio, significantly outperforming CPU-based setups. Llama 3.2 provides the best balance of accuracy and power efficiency, while TinyLlama is well-suited for low-power environments at the cost of reduced accuracy. In contrast, Phi-3 Mini consumes the most energy despite its high accuracy. In addition, GPU acceleration, memory bandwidth, and model architecture are key in optimizing inference energy efficiency. Our empirical analysis offers practical insights for AI, smart systems, and mobile ad-hoc platforms to leverage tradeoffs from accuracy, inference latency, and power efficiency in energy-constrained environments.
title Characterizing and Understanding Energy Footprint and Efficiency of Small Language Model on Edges
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2511.11624