BitNet Distillation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wu, Xun, Huang, Shaohan, Wang, Wenhui, Song, Ting, Dong, Li, Xia, Yan, Wei, Furu
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908596622589952
author Wu, Xun
Huang, Shaohan
Wang, Wenhui
Song, Ting
Dong, Li
Xia, Yan
Wei, Furu
author_facet Wu, Xun
Huang, Shaohan
Wang, Wenhui
Song, Ting
Dong, Li
Xia, Yan
Wei, Furu
contents In this paper, we present BitNet Distillation (BitDistill), a lightweight pipeline that fine-tunes off-the-shelf full-precision LLMs (e.g., Qwen) into 1.58-bit precision (i.e., ternary weights {-1, 0, 1}) for specific downstream tasks, achieving strong task-specific performance with minimal computational cost. Specifically, BitDistill incorporates three key techniques: the SubLN module, as introduced in BitNet; multi-head attention distillation, based on MiniLM; and continual pre-training, which serves as a crucial warm-up step to mitigate the scalability issue of the performance gap between finetuned full-precision and 1.58-bit LLMs on specific tasks. Experimental results show that BitDistill achieves performance comparable to the full-precision counterpart models across model size, while enabling up to 10x memory savings and 2.65x faster inference on CPUs. Code is available at https://github.com/microsoft/BitNet.
format Preprint
id arxiv_https___arxiv_org_abs_2510_13998
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BitNet Distillation
Wu, Xun
Huang, Shaohan
Wang, Wenhui
Song, Ting
Dong, Li
Xia, Yan
Wei, Furu
Machine Learning
Computation and Language
In this paper, we present BitNet Distillation (BitDistill), a lightweight pipeline that fine-tunes off-the-shelf full-precision LLMs (e.g., Qwen) into 1.58-bit precision (i.e., ternary weights {-1, 0, 1}) for specific downstream tasks, achieving strong task-specific performance with minimal computational cost. Specifically, BitDistill incorporates three key techniques: the SubLN module, as introduced in BitNet; multi-head attention distillation, based on MiniLM; and continual pre-training, which serves as a crucial warm-up step to mitigate the scalability issue of the performance gap between finetuned full-precision and 1.58-bit LLMs on specific tasks. Experimental results show that BitDistill achieves performance comparable to the full-precision counterpart models across model size, while enabling up to 10x memory savings and 2.65x faster inference on CPUs. Code is available at https://github.com/microsoft/BitNet.
title BitNet Distillation
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2510.13998