TinyLLM: A Framework for Training and Deploying Language Models at the Edge Computers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kandala, Savitha Viswanadh, Medaranga, Pramuka, Varshney, Ambuj
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912162640822272
author Kandala, Savitha Viswanadh
Medaranga, Pramuka
Varshney, Ambuj
author_facet Kandala, Savitha Viswanadh
Medaranga, Pramuka
Varshney, Ambuj
contents Language models have gained significant interest due to their general-purpose capabilities, which appear to emerge as models are scaled to increasingly larger parameter sizes. However, these large models impose stringent requirements on computing systems, necessitating significant memory and processing requirements for inference. This makes performing inference on mobile and edge devices challenging, often requiring invocating remotely-hosted models via network calls. Remote inference, in turn, introduces issues like latency, unreliable network connectivity, and privacy concerns. To address these challenges, we explored the possibility of deviating from the trend of increasing model size. Instead, we hypothesize that much smaller models (~30-120M parameters) can outperform their larger counterparts for specific tasks by carefully curating the data used for pre-training and fine-tuning. We investigate this within the context of deploying edge-device models to support sensing applications. We trained several foundational models through a systematic study and found that small models can run locally on edge devices, achieving high token rates and accuracy. Based on these findings, we developed a framework that allows users to train foundational models tailored to their specific applications and deploy them at the edge.
format Preprint
id arxiv_https___arxiv_org_abs_2412_15304
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TinyLLM: A Framework for Training and Deploying Language Models at the Edge Computers
Kandala, Savitha Viswanadh
Medaranga, Pramuka
Varshney, Ambuj
Machine Learning
Distributed, Parallel, and Cluster Computing
Emerging Technologies
Networking and Internet Architecture
Language models have gained significant interest due to their general-purpose capabilities, which appear to emerge as models are scaled to increasingly larger parameter sizes. However, these large models impose stringent requirements on computing systems, necessitating significant memory and processing requirements for inference. This makes performing inference on mobile and edge devices challenging, often requiring invocating remotely-hosted models via network calls. Remote inference, in turn, introduces issues like latency, unreliable network connectivity, and privacy concerns. To address these challenges, we explored the possibility of deviating from the trend of increasing model size. Instead, we hypothesize that much smaller models (~30-120M parameters) can outperform their larger counterparts for specific tasks by carefully curating the data used for pre-training and fine-tuning. We investigate this within the context of deploying edge-device models to support sensing applications. We trained several foundational models through a systematic study and found that small models can run locally on edge devices, achieving high token rates and accuracy. Based on these findings, we developed a framework that allows users to train foundational models tailored to their specific applications and deploy them at the edge.
title TinyLLM: A Framework for Training and Deploying Language Models at the Edge Computers
topic Machine Learning
Distributed, Parallel, and Cluster Computing
Emerging Technologies
Networking and Internet Architecture
url https://arxiv.org/abs/2412.15304