MobiEdit: Resource-efficient Knowledge Editing for Personalized On-device LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Zhenyan, Xu, Daliang, Cai, Dongqi, Li, Zexi, Liu, Wei, Liu, Fangming, Wang, Shangguang, Xu, Mengwei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911007666864128
author Lu, Zhenyan
Xu, Daliang
Cai, Dongqi
Li, Zexi
Liu, Wei
Liu, Fangming
Wang, Shangguang
Xu, Mengwei
author_facet Lu, Zhenyan
Xu, Daliang
Cai, Dongqi
Li, Zexi
Liu, Wei
Liu, Fangming
Wang, Shangguang
Xu, Mengwei
contents Large language models (LLMs) are deployed on mobile devices to power killer applications such as intelligent assistants. LLMs pre-trained on general corpora often hallucinate when handling personalized or unseen queries, leading to incorrect or outdated responses. Knowledge editing addresses this by identifying and adjusting a small crucial portion of model weights, without compromising the general knowledge. However, prior knowledge editing methods are impractical to run on local devices due to the resource-heavy backpropagation (BP) needed for updates. We present MobiEdit, the first mobile knowledge editing framework that enables efficient LLM personalization on commercial off-the-shelf (COTS) mobile devices. MobiEdit replaces full-precision BP with quantized forward-only gradient estimation, thus compatible with the energy-efficient mobile neural processing units (NPUs). MobiEdit replaces full-precision backpropagation with quantized forward-only gradient estimation, making it compatible with energy-efficient mobile NPUs. To further improve gradient estimation efficiency, we introduce two optimizations: an early stoping mechanism that adaptively terminates editing upon success and a prefix cache that reuses computation across steps. Our approach enables real-time editing of a 3B-parameter model (Qwen2.5-3B-Instruct) on COTS mobile devices with 7.6$\times$ less memory, 14.7 $\times$ less energy and 3.6$\times$ less latency compared to previous knowledge editing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13772
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MobiEdit: Resource-efficient Knowledge Editing for Personalized On-device LLMs
Lu, Zhenyan
Xu, Daliang
Cai, Dongqi
Li, Zexi
Liu, Wei
Liu, Fangming
Wang, Shangguang
Xu, Mengwei
Machine Learning
Artificial Intelligence
Large language models (LLMs) are deployed on mobile devices to power killer applications such as intelligent assistants. LLMs pre-trained on general corpora often hallucinate when handling personalized or unseen queries, leading to incorrect or outdated responses. Knowledge editing addresses this by identifying and adjusting a small crucial portion of model weights, without compromising the general knowledge. However, prior knowledge editing methods are impractical to run on local devices due to the resource-heavy backpropagation (BP) needed for updates. We present MobiEdit, the first mobile knowledge editing framework that enables efficient LLM personalization on commercial off-the-shelf (COTS) mobile devices. MobiEdit replaces full-precision BP with quantized forward-only gradient estimation, thus compatible with the energy-efficient mobile neural processing units (NPUs). MobiEdit replaces full-precision backpropagation with quantized forward-only gradient estimation, making it compatible with energy-efficient mobile NPUs. To further improve gradient estimation efficiency, we introduce two optimizations: an early stoping mechanism that adaptively terminates editing upon success and a prefix cache that reuses computation across steps. Our approach enables real-time editing of a 3B-parameter model (Qwen2.5-3B-Instruct) on COTS mobile devices with 7.6$\times$ less memory, 14.7 $\times$ less energy and 3.6$\times$ less latency compared to previous knowledge editing methods.
title MobiEdit: Resource-efficient Knowledge Editing for Personalized On-device LLMs
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.13772