Computational Efficiency of Large Language Models (LLMs) in Resource-Constrained Environments

Fuente: Zenodo
Gespeichert in:
Bibliographische Detailangaben
1. Verfasser: Mission Franklin
Format: Recurso digital
Sprache:Englisch
Veröffentlicht: Zenodo 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866902073264570368
author Mission Franklin
author_facet Mission Franklin
contents <p><em><span>Large Language Models (LLMs) such as GPT-3 and BERT have significantly improved the performance of natural language processing applications, including chatbots, machine translation, content generation, and virtual assistants. Despite their high accuracy and advanced language understanding capabilities, these models require substantial computational resources, including powerful GPUs, large memory capacity, and high energy consumption. Such requirements make the deployment of LLMs difficult in resource-constrained environments such as mobile devices, Internet of Things (IoT) systems, embedded platforms, and edge computing infrastructures.</span></em></p> <p><em><span>This study focuses on improving the computational efficiency of Large Language Models while maintaining acceptable performance levels. The research examines the major challenges associated with deploying LLMs in low-resource environments and reviews common optimization techniques such as model compression, pruning, quantization, knowledge distillation, and parameter-efficient fine-tuning. The study also explores the trade-off between model accuracy and computational efficiency and highlights the importance of lightweight and scalable AI solutions for edge computing applications.</span></em></p> <p><em><span>The findings suggest that optimization methods can significantly reduce model size, inference time, memory usage, and energy consumption, making LLMs more practical for real-world deployment on low-power devices. However, balancing efficiency and performance remains a major challenge, as excessive reduction in computational requirements may negatively affect model accuracy. The study concludes by recommending practical strategies for developing efficient, accessible, and scalable LLMs suitable for diverse resource-constrained environments.</span></em></p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_20321300
institution Zenodo
language eng
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle Computational Efficiency of Large Language Models (LLMs) in Resource-Constrained Environments
Mission Franklin
Large Language Models (LLMs), Computational Efficiency, Resource-Constrained Environments, Model Compression, Quantization, Pruning, Knowledge Distillation, Parameter-Efficient Fine-Tuning, Edge Computing, Artificial Intelligence.
<p><em><span>Large Language Models (LLMs) such as GPT-3 and BERT have significantly improved the performance of natural language processing applications, including chatbots, machine translation, content generation, and virtual assistants. Despite their high accuracy and advanced language understanding capabilities, these models require substantial computational resources, including powerful GPUs, large memory capacity, and high energy consumption. Such requirements make the deployment of LLMs difficult in resource-constrained environments such as mobile devices, Internet of Things (IoT) systems, embedded platforms, and edge computing infrastructures.</span></em></p> <p><em><span>This study focuses on improving the computational efficiency of Large Language Models while maintaining acceptable performance levels. The research examines the major challenges associated with deploying LLMs in low-resource environments and reviews common optimization techniques such as model compression, pruning, quantization, knowledge distillation, and parameter-efficient fine-tuning. The study also explores the trade-off between model accuracy and computational efficiency and highlights the importance of lightweight and scalable AI solutions for edge computing applications.</span></em></p> <p><em><span>The findings suggest that optimization methods can significantly reduce model size, inference time, memory usage, and energy consumption, making LLMs more practical for real-world deployment on low-power devices. However, balancing efficiency and performance remains a major challenge, as excessive reduction in computational requirements may negatively affect model accuracy. The study concludes by recommending practical strategies for developing efficient, accessible, and scalable LLMs suitable for diverse resource-constrained environments.</span></em></p>
title Computational Efficiency of Large Language Models (LLMs) in Resource-Constrained Environments
topic Large Language Models (LLMs), Computational Efficiency, Resource-Constrained Environments, Model Compression, Quantization, Pruning, Knowledge Distillation, Parameter-Efficient Fine-Tuning, Edge Computing, Artificial Intelligence.
url https://doi.org/10.5281/zenodo.20321300