TQCompressor: improving tensor decomposition methods in neural networks via permutations

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Abronin, V., Naumov, A., Mazur, D., Bystrov, D., Tsarova, K., Melnikov, Ar., Oseledets, I., Dolgov, S., Brasher, R., Perelshtein, M.
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913213989257216
author Abronin, V.
Naumov, A.
Mazur, D.
Bystrov, D.
Tsarova, K.
Melnikov, Ar.
Oseledets, I.
Dolgov, S.
Brasher, R.
Perelshtein, M.
author_facet Abronin, V.
Naumov, A.
Mazur, D.
Bystrov, D.
Tsarova, K.
Melnikov, Ar.
Oseledets, I.
Dolgov, S.
Brasher, R.
Perelshtein, M.
contents We introduce TQCompressor, a novel method for neural network model compression with improved tensor decompositions. We explore the challenges posed by the computational and storage demands of pre-trained language models in NLP tasks and propose a permutation-based enhancement to Kronecker decomposition. This enhancement makes it possible to reduce loss in model expressivity which is usually associated with factorization. We demonstrate this method applied to the GPT-2$_{small}$. The result of the compression is TQCompressedGPT-2 model, featuring 81 mln. parameters compared to 124 mln. in the GPT-2$_{small}$. We make TQCompressedGPT-2 publicly available. We further enhance the performance of the TQCompressedGPT-2 through a training strategy involving multi-step knowledge distillation, using only a 3.1% of the OpenWebText. TQCompressedGPT-2 surpasses DistilGPT-2 and KnGPT-2 in comparative evaluations, marking an advancement in the efficient and effective deployment of models in resource-constrained environments.
format Preprint
id arxiv_https___arxiv_org_abs_2401_16367
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TQCompressor: improving tensor decomposition methods in neural networks via permutations
Abronin, V.
Naumov, A.
Mazur, D.
Bystrov, D.
Tsarova, K.
Melnikov, Ar.
Oseledets, I.
Dolgov, S.
Brasher, R.
Perelshtein, M.
Machine Learning
Artificial Intelligence
Computation and Language
We introduce TQCompressor, a novel method for neural network model compression with improved tensor decompositions. We explore the challenges posed by the computational and storage demands of pre-trained language models in NLP tasks and propose a permutation-based enhancement to Kronecker decomposition. This enhancement makes it possible to reduce loss in model expressivity which is usually associated with factorization. We demonstrate this method applied to the GPT-2$_{small}$. The result of the compression is TQCompressedGPT-2 model, featuring 81 mln. parameters compared to 124 mln. in the GPT-2$_{small}$. We make TQCompressedGPT-2 publicly available. We further enhance the performance of the TQCompressedGPT-2 through a training strategy involving multi-step knowledge distillation, using only a 3.1% of the OpenWebText. TQCompressedGPT-2 surpasses DistilGPT-2 and KnGPT-2 in comparative evaluations, marking an advancement in the efficient and effective deployment of models in resource-constrained environments.
title TQCompressor: improving tensor decomposition methods in neural networks via permutations
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2401.16367