Projected Compression: Trainable Projection for Efficient Transformer Compression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Stefaniak, Maciej, Krutul, Michał, Małaśnicki, Jan, Pióro, Maciej, Krajewski, Jakub, Jaszczur, Sebastian, Cygan, Marek, Adamczewski, Kamil, Ludziejewski, Jan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918073387188224
author Stefaniak, Maciej
Krutul, Michał
Małaśnicki, Jan
Pióro, Maciej
Krajewski, Jakub
Jaszczur, Sebastian
Cygan, Marek
Adamczewski, Kamil
Ludziejewski, Jan
author_facet Stefaniak, Maciej
Krutul, Michał
Małaśnicki, Jan
Pióro, Maciej
Krajewski, Jakub
Jaszczur, Sebastian
Cygan, Marek
Adamczewski, Kamil
Ludziejewski, Jan
contents Large language models have steadily increased in size to achieve improved performance; however, this growth has also led to greater inference time and computational demands. Consequently, there is rising interest in model size reduction methods. To address this issue, we propose Projected Compression, a novel model compression technique, that reduces model weights by utilizing projection modules. Specifically, we first train additional trainable projections weights and preserve access to all the original model parameters. Subsequently, these projections are merged into a lower-dimensional product matrix, resulting in a reduced-size standard Transformer-based model. Unlike alternative approaches that require additional computational overhead, our method matches the base model's per-token computation step in FLOPs. Experimental results show that Projected Compression outperforms the comparable hard pruning and retraining approach on higher quality models. Moreover, the performance margin scales well with the number of tokens.
format Preprint
id arxiv_https___arxiv_org_abs_2506_22255
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Projected Compression: Trainable Projection for Efficient Transformer Compression
Stefaniak, Maciej
Krutul, Michał
Małaśnicki, Jan
Pióro, Maciej
Krajewski, Jakub
Jaszczur, Sebastian
Cygan, Marek
Adamczewski, Kamil
Ludziejewski, Jan
Machine Learning
Artificial Intelligence
Computation and Language
Large language models have steadily increased in size to achieve improved performance; however, this growth has also led to greater inference time and computational demands. Consequently, there is rising interest in model size reduction methods. To address this issue, we propose Projected Compression, a novel model compression technique, that reduces model weights by utilizing projection modules. Specifically, we first train additional trainable projections weights and preserve access to all the original model parameters. Subsequently, these projections are merged into a lower-dimensional product matrix, resulting in a reduced-size standard Transformer-based model. Unlike alternative approaches that require additional computational overhead, our method matches the base model's per-token computation step in FLOPs. Experimental results show that Projected Compression outperforms the comparable hard pruning and retraining approach on higher quality models. Moreover, the performance margin scales well with the number of tokens.
title Projected Compression: Trainable Projection for Efficient Transformer Compression
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2506.22255