Krony-PT: GPT2 compressed with Kronecker Products
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914068713963520 |
|---|---|
| author | Ayad, Mohamed Ayoub Ben Mitrovic, Jelena Granitzer, Michael |
| author_facet | Ayad, Mohamed Ayoub Ben Mitrovic, Jelena Granitzer, Michael |
| contents | We introduce Krony-PT, a compression technique for GPT-2 based on Kronecker products. We specifically target the feed-forward weights of each transformer block, and systematically compress the feed-forward layer matrices to various degrees. We introduce a modified Van Loan decomposition to initialize new Kronecker factors, and also propose a new pruning-based initialization technique. Our method compresses the original 124M-parameter GPT-2 to various smaller models, ranging from 80M to 96M. Our 81M model variant outperforms DistilGPT2 on next-token prediction across all standard language modeling datasets, and shows competitive or comparable performance with significantly larger Kronecker-based compressions of GPT-2. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_12351 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Krony-PT: GPT2 compressed with Kronecker Products Ayad, Mohamed Ayoub Ben Mitrovic, Jelena Granitzer, Michael Machine Learning Computation and Language We introduce Krony-PT, a compression technique for GPT-2 based on Kronecker products. We specifically target the feed-forward weights of each transformer block, and systematically compress the feed-forward layer matrices to various degrees. We introduce a modified Van Loan decomposition to initialize new Kronecker factors, and also propose a new pruning-based initialization technique. Our method compresses the original 124M-parameter GPT-2 to various smaller models, ranging from 80M to 96M. Our 81M model variant outperforms DistilGPT2 on next-token prediction across all standard language modeling datasets, and shows competitive or comparable performance with significantly larger Kronecker-based compressions of GPT-2. |
| title | Krony-PT: GPT2 compressed with Kronecker Products |
| topic | Machine Learning Computation and Language |
| url | https://arxiv.org/abs/2412.12351 |