Hyper-Compression: Model Compression via Hyperfunction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fan, Fenglei, Fan, Juntong, Wang, Dayang, Zhang, Jingbo, Dong, Zelin, Zhang, Shijun, Wang, Ge, Zeng, Tieyong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915769018744832
author Fan, Fenglei
Fan, Juntong
Wang, Dayang
Zhang, Jingbo
Dong, Zelin
Zhang, Shijun
Wang, Ge
Zeng, Tieyong
author_facet Fan, Fenglei
Fan, Juntong
Wang, Dayang
Zhang, Jingbo
Dong, Zelin
Zhang, Shijun
Wang, Ge
Zeng, Tieyong
contents The rapid growth of large models' size has far outpaced that of computing resources. To bridge this gap, encouraged by the parsimonious relationship between genotype and phenotype in the brain's growth and development, we propose the so-called Hyper-Compression that turns the model compression into the issue of parameter representation via a hyperfunction. Specifically, it is known that the trajectory of some low-dimensional dynamic systems can fill the high-dimensional space eventually. Thus, Hyper-Compression, using these dynamic systems as the hyperfunctions, represents the parameters of the target network by their corresponding composition number or trajectory length. This suggests a novel mechanism for model compression, substantially different from the existing pruning, quantization, distillation, and decomposition. Along this direction, we methodologically identify a suitable dynamic system with the irrational winding as the hyperfunction and theoretically derive its associated error bound. Next, guided by our theoretical insights, we propose several engineering twists to make the Hyper-Compression pragmatic and effective. Lastly, systematic and comprehensive experiments on \textcolor{black}{NLP models such as LLaMA and Qwen series and vision models} confirm that Hyper-Compression enjoys the following \textbf{PNAS} merits: 1) \textbf{P}referable compression ratio; 2) \textbf{N}o post-hoc retraining; 3) \textbf{A}ffordable inference time; and 4) \textbf{S}hort compression time. It compresses LLaMA2-7B in an hour and achieves close-to-int4-quantization performance, without retraining and with a performance drop of less than 1\%. We have open-sourced our code in https://github.com/Juntongkuki/Hyper-Compression.git for free download and evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2409_00592
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Hyper-Compression: Model Compression via Hyperfunction
Fan, Fenglei
Fan, Juntong
Wang, Dayang
Zhang, Jingbo
Dong, Zelin
Zhang, Shijun
Wang, Ge
Zeng, Tieyong
Machine Learning
Artificial Intelligence
Emerging Technologies
The rapid growth of large models' size has far outpaced that of computing resources. To bridge this gap, encouraged by the parsimonious relationship between genotype and phenotype in the brain's growth and development, we propose the so-called Hyper-Compression that turns the model compression into the issue of parameter representation via a hyperfunction. Specifically, it is known that the trajectory of some low-dimensional dynamic systems can fill the high-dimensional space eventually. Thus, Hyper-Compression, using these dynamic systems as the hyperfunctions, represents the parameters of the target network by their corresponding composition number or trajectory length. This suggests a novel mechanism for model compression, substantially different from the existing pruning, quantization, distillation, and decomposition. Along this direction, we methodologically identify a suitable dynamic system with the irrational winding as the hyperfunction and theoretically derive its associated error bound. Next, guided by our theoretical insights, we propose several engineering twists to make the Hyper-Compression pragmatic and effective. Lastly, systematic and comprehensive experiments on \textcolor{black}{NLP models such as LLaMA and Qwen series and vision models} confirm that Hyper-Compression enjoys the following \textbf{PNAS} merits: 1) \textbf{P}referable compression ratio; 2) \textbf{N}o post-hoc retraining; 3) \textbf{A}ffordable inference time; and 4) \textbf{S}hort compression time. It compresses LLaMA2-7B in an hour and achieves close-to-int4-quantization performance, without retraining and with a performance drop of less than 1\%. We have open-sourced our code in https://github.com/Juntongkuki/Hyper-Compression.git for free download and evaluation.
title Hyper-Compression: Model Compression via Hyperfunction
topic Machine Learning
Artificial Intelligence
Emerging Technologies
url https://arxiv.org/abs/2409.00592