SlimLM: An Efficient Small Language Model for On-Device Document Assistance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pham, Thang M., Nguyen, Phat T., Yoon, Seunghyun, Lai, Viet Dac, Dernoncourt, Franck, Bui, Trung
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929604177952768
author Pham, Thang M.
Nguyen, Phat T.
Yoon, Seunghyun
Lai, Viet Dac
Dernoncourt, Franck
Bui, Trung
author_facet Pham, Thang M.
Nguyen, Phat T.
Yoon, Seunghyun
Lai, Viet Dac
Dernoncourt, Franck
Bui, Trung
contents While small language models (SLMs) show promises for mobile deployment, their real-world performance and applications on smartphones remains underexplored. We present SlimLM, a series of SLMs optimized for document assistance tasks on mobile devices. Through extensive experiments on a Samsung Galaxy S24, we identify the optimal trade-offs between model size (ranging from 125M to 7B parameters), context length, and inference time for efficient on-device processing. SlimLM is pre-trained on SlimPajama-627B and fine-tuned on DocAssist, our constructed dataset for summarization, question answering and suggestion tasks. Our smallest model demonstrates efficient performance on S24, while larger variants offer enhanced capabilities within mobile constraints. We evaluate SlimLM against existing SLMs, showing comparable or superior performance and offering a benchmark for future research in on-device language models. We also provide an Android application, offering practical insights into SLM deployment. Our findings provide valuable insights and illuminate the capabilities of running advanced language models on high-end smartphones, potentially reducing server costs and enhancing privacy through on-device processing.
format Preprint
id arxiv_https___arxiv_org_abs_2411_09944
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SlimLM: An Efficient Small Language Model for On-Device Document Assistance
Pham, Thang M.
Nguyen, Phat T.
Yoon, Seunghyun
Lai, Viet Dac
Dernoncourt, Franck
Bui, Trung
Computation and Language
While small language models (SLMs) show promises for mobile deployment, their real-world performance and applications on smartphones remains underexplored. We present SlimLM, a series of SLMs optimized for document assistance tasks on mobile devices. Through extensive experiments on a Samsung Galaxy S24, we identify the optimal trade-offs between model size (ranging from 125M to 7B parameters), context length, and inference time for efficient on-device processing. SlimLM is pre-trained on SlimPajama-627B and fine-tuned on DocAssist, our constructed dataset for summarization, question answering and suggestion tasks. Our smallest model demonstrates efficient performance on S24, while larger variants offer enhanced capabilities within mobile constraints. We evaluate SlimLM against existing SLMs, showing comparable or superior performance and offering a benchmark for future research in on-device language models. We also provide an Android application, offering practical insights into SLM deployment. Our findings provide valuable insights and illuminate the capabilities of running advanced language models on high-end smartphones, potentially reducing server costs and enhancing privacy through on-device processing.
title SlimLM: An Efficient Small Language Model for On-Device Document Assistance
topic Computation and Language
url https://arxiv.org/abs/2411.09944