Thanos: A Block-wise Pruning Algorithm for Efficient Large Language Model Compression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ilin, Ivan, Richtarik, Peter
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908305942642688
author Ilin, Ivan
Richtarik, Peter
author_facet Ilin, Ivan
Richtarik, Peter
contents This paper presents Thanos, a novel weight-pruning algorithm designed to reduce the memory footprint and enhance the computational efficiency of large language models (LLMs) by removing redundant weights while maintaining accuracy. Thanos introduces a block-wise pruning strategy with adaptive masks that dynamically adjust to weight importance, enabling flexible sparsity patterns and structured formats, such as $n:m$ sparsity, optimized for hardware acceleration. Experimental evaluations demonstrate that Thanos achieves state-of-the-art performance in structured pruning and outperforms existing methods in unstructured pruning. By providing an efficient and adaptable approach to model compression, Thanos offers a practical solution for deploying large models in resource-constrained environments.
format Preprint
id arxiv_https___arxiv_org_abs_2504_05346
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Thanos: A Block-wise Pruning Algorithm for Efficient Large Language Model Compression
Ilin, Ivan
Richtarik, Peter
Machine Learning
Artificial Intelligence
Computation and Language
Performance
68T07, 68Q32
This paper presents Thanos, a novel weight-pruning algorithm designed to reduce the memory footprint and enhance the computational efficiency of large language models (LLMs) by removing redundant weights while maintaining accuracy. Thanos introduces a block-wise pruning strategy with adaptive masks that dynamically adjust to weight importance, enabling flexible sparsity patterns and structured formats, such as $n:m$ sparsity, optimized for hardware acceleration. Experimental evaluations demonstrate that Thanos achieves state-of-the-art performance in structured pruning and outperforms existing methods in unstructured pruning. By providing an efficient and adaptable approach to model compression, Thanos offers a practical solution for deploying large models in resource-constrained environments.
title Thanos: A Block-wise Pruning Algorithm for Efficient Large Language Model Compression
topic Machine Learning
Artificial Intelligence
Computation and Language
Performance
68T07, 68Q32
url https://arxiv.org/abs/2504.05346