PocketLLM: Ultimate Compression of Large Language Models via Meta Networks

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Tian, Ye, Wang, Chengcheng, Han, Jing, Tang, Yehui, Han, Kai
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908669396910080
author Tian, Ye
Wang, Chengcheng
Han, Jing
Tang, Yehui
Han, Kai
author_facet Tian, Ye
Wang, Chengcheng
Han, Jing
Tang, Yehui
Han, Kai
contents As Large Language Models (LLMs) continue to grow in size, storing and transmitting them on edge devices becomes increasingly challenging. Traditional methods like quantization and pruning struggle to achieve extreme compression of LLMs without sacrificing accuracy. In this paper, we introduce PocketLLM, a novel approach to compress LLMs in a latent space via meta-networks. A simple encoder network is proposed to project the weights of LLMs into discrete latent vectors, which are then represented using a compact codebook. A lightweight decoder network is employed to map the codebook's representative vectors back to the original weight space. This method allows for significant compression of the large weights in LLMs, consisting solely of a small decoder, a concise codebook, and an index. Extensive experiments show that PocketLLM achieves superior performance even at significantly high compression ratios, e.g., compressing Llama 2-7B by 10x with a negligible drop in accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2511_17637
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PocketLLM: Ultimate Compression of Large Language Models via Meta Networks
Tian, Ye
Wang, Chengcheng
Han, Jing
Tang, Yehui
Han, Kai
Machine Learning
Computation and Language
As Large Language Models (LLMs) continue to grow in size, storing and transmitting them on edge devices becomes increasingly challenging. Traditional methods like quantization and pruning struggle to achieve extreme compression of LLMs without sacrificing accuracy. In this paper, we introduce PocketLLM, a novel approach to compress LLMs in a latent space via meta-networks. A simple encoder network is proposed to project the weights of LLMs into discrete latent vectors, which are then represented using a compact codebook. A lightweight decoder network is employed to map the codebook's representative vectors back to the original weight space. This method allows for significant compression of the large weights in LLMs, consisting solely of a small decoder, a concise codebook, and an index. Extensive experiments show that PocketLLM achieves superior performance even at significantly high compression ratios, e.g., compressing Llama 2-7B by 10x with a negligible drop in accuracy.
title PocketLLM: Ultimate Compression of Large Language Models via Meta Networks
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2511.17637