Uncertainty-Aware Hybrid Inference with On-Device Small and Remote Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Oh, Seungeun, Kim, Jinhyuk, Park, Jihong, Ko, Seung-Woo, Quek, Tony Q. S., Kim, Seong-Lyun
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915203375955968
author Oh, Seungeun
Kim, Jinhyuk
Park, Jihong
Ko, Seung-Woo
Quek, Tony Q. S.
Kim, Seong-Lyun
author_facet Oh, Seungeun
Kim, Jinhyuk
Park, Jihong
Ko, Seung-Woo
Quek, Tony Q. S.
Kim, Seong-Lyun
contents This paper studies a hybrid language model (HLM) architecture that integrates a small language model (SLM) operating on a mobile device with a large language model (LLM) hosted at the base station (BS) of a wireless network. The HLM token generation process follows the speculative inference principle: the SLM's vocabulary distribution is uploaded to the LLM, which either accepts or rejects it, with rejected tokens being resampled by the LLM. While this approach ensures alignment between the vocabulary distributions of the SLM and LLM, it suffers from low token throughput due to uplink transmission and the computation costs of running both language models. To address this, we propose a novel HLM structure coined Uncertainty-aware opportunistic HLM (U-HLM), wherein the SLM locally measures its output uncertainty and skips both uplink transmissions and LLM operations for tokens that are likely to be accepted. This opportunistic skipping is enabled by our empirical finding of a linear correlation between the SLM's uncertainty and the LLM's rejection probability. We analytically derive the uncertainty threshold and evaluate its expected risk of rejection. Simulations show that U-HLM reduces uplink transmissions and LLM computations by 45.93%, while achieving up to 97.54% of the LLM's inference accuracy and 2.54$\times$ faster token throughput than HLM without skipping.
format Preprint
id arxiv_https___arxiv_org_abs_2412_12687
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Uncertainty-Aware Hybrid Inference with On-Device Small and Remote Large Language Models
Oh, Seungeun
Kim, Jinhyuk
Park, Jihong
Ko, Seung-Woo
Quek, Tony Q. S.
Kim, Seong-Lyun
Machine Learning
Distributed, Parallel, and Cluster Computing
Information Theory
Networking and Internet Architecture
Signal Processing
This paper studies a hybrid language model (HLM) architecture that integrates a small language model (SLM) operating on a mobile device with a large language model (LLM) hosted at the base station (BS) of a wireless network. The HLM token generation process follows the speculative inference principle: the SLM's vocabulary distribution is uploaded to the LLM, which either accepts or rejects it, with rejected tokens being resampled by the LLM. While this approach ensures alignment between the vocabulary distributions of the SLM and LLM, it suffers from low token throughput due to uplink transmission and the computation costs of running both language models. To address this, we propose a novel HLM structure coined Uncertainty-aware opportunistic HLM (U-HLM), wherein the SLM locally measures its output uncertainty and skips both uplink transmissions and LLM operations for tokens that are likely to be accepted. This opportunistic skipping is enabled by our empirical finding of a linear correlation between the SLM's uncertainty and the LLM's rejection probability. We analytically derive the uncertainty threshold and evaluate its expected risk of rejection. Simulations show that U-HLM reduces uplink transmissions and LLM computations by 45.93%, while achieving up to 97.54% of the LLM's inference accuracy and 2.54$\times$ faster token throughput than HLM without skipping.
title Uncertainty-Aware Hybrid Inference with On-Device Small and Remote Large Language Models
topic Machine Learning
Distributed, Parallel, and Cluster Computing
Information Theory
Networking and Internet Architecture
Signal Processing
url https://arxiv.org/abs/2412.12687