Thanos: A Block-wise Pruning Algorithm for Efficient Large Language Model Compression
Fuente:
arXiv
Saved in:
| Main Authors: | , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908305942642688 |
|---|---|
| author | Ilin, Ivan Richtarik, Peter |
| author_facet | Ilin, Ivan Richtarik, Peter |
| contents | This paper presents Thanos, a novel weight-pruning algorithm designed to reduce the memory footprint and enhance the computational efficiency of large language models (LLMs) by removing redundant weights while maintaining accuracy. Thanos introduces a block-wise pruning strategy with adaptive masks that dynamically adjust to weight importance, enabling flexible sparsity patterns and structured formats, such as $n:m$ sparsity, optimized for hardware acceleration. Experimental evaluations demonstrate that Thanos achieves state-of-the-art performance in structured pruning and outperforms existing methods in unstructured pruning. By providing an efficient and adaptable approach to model compression, Thanos offers a practical solution for deploying large models in resource-constrained environments. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_05346 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Thanos: A Block-wise Pruning Algorithm for Efficient Large Language Model Compression Ilin, Ivan Richtarik, Peter Machine Learning Artificial Intelligence Computation and Language Performance 68T07, 68Q32 This paper presents Thanos, a novel weight-pruning algorithm designed to reduce the memory footprint and enhance the computational efficiency of large language models (LLMs) by removing redundant weights while maintaining accuracy. Thanos introduces a block-wise pruning strategy with adaptive masks that dynamically adjust to weight importance, enabling flexible sparsity patterns and structured formats, such as $n:m$ sparsity, optimized for hardware acceleration. Experimental evaluations demonstrate that Thanos achieves state-of-the-art performance in structured pruning and outperforms existing methods in unstructured pruning. By providing an efficient and adaptable approach to model compression, Thanos offers a practical solution for deploying large models in resource-constrained environments. |
| title | Thanos: A Block-wise Pruning Algorithm for Efficient Large Language Model Compression |
| topic | Machine Learning Artificial Intelligence Computation and Language Performance 68T07, 68Q32 |
| url | https://arxiv.org/abs/2504.05346 |