Efficient Mathematical Reasoning Models via Dynamic Pruning and Knowledge Distillation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yu, Fengming, Meng, Qingyu, Pan, Haiwei, Zhang, Kejia
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917098093019136
author Yu, Fengming
Meng, Qingyu
Pan, Haiwei
Zhang, Kejia
author_facet Yu, Fengming
Meng, Qingyu
Pan, Haiwei
Zhang, Kejia
contents With the rapid development of deep learning, large language models have shown strong capabilities in complex reasoning tasks such as mathematical equation solving. However, their substantial computational and storage costs hinder practical deployment. This paper proposes a lightweight optimization method that integrates dynamic attention head pruning with knowledge distillation. The approach dynamically evaluates the importance of each attention head in the multi-head attention mechanism using a combination of weight norms and entropy, and prunes redundant heads in real time to reduce computational overhead. To mitigate performance degradation, knowledge distillation transfers information from the original model to the pruned student, enabling the smaller model to preserve reasoning ability. Experiments conducted on both Math23k and ASDiv-A verify the effectiveness of the proposed method. For example, on Math23k with a 30% pruning ratio, parameters are reduced by 18.7%, inference speed is improved by 27.5%, FLOPs are reduced by 19.3%, and accuracy drops only 0.7% (from 84.4% to 83.7%). These results demonstrate that the method achieves substantial efficiency gains while maintaining strong reasoning performance, providing a practical solution for efficient deployment of large language models in mathematical reasoning tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2511_17577
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Mathematical Reasoning Models via Dynamic Pruning and Knowledge Distillation
Yu, Fengming
Meng, Qingyu
Pan, Haiwei
Zhang, Kejia
Machine Learning
Artificial Intelligence
Computation and Language
With the rapid development of deep learning, large language models have shown strong capabilities in complex reasoning tasks such as mathematical equation solving. However, their substantial computational and storage costs hinder practical deployment. This paper proposes a lightweight optimization method that integrates dynamic attention head pruning with knowledge distillation. The approach dynamically evaluates the importance of each attention head in the multi-head attention mechanism using a combination of weight norms and entropy, and prunes redundant heads in real time to reduce computational overhead. To mitigate performance degradation, knowledge distillation transfers information from the original model to the pruned student, enabling the smaller model to preserve reasoning ability. Experiments conducted on both Math23k and ASDiv-A verify the effectiveness of the proposed method. For example, on Math23k with a 30% pruning ratio, parameters are reduced by 18.7%, inference speed is improved by 27.5%, FLOPs are reduced by 19.3%, and accuracy drops only 0.7% (from 84.4% to 83.7%). These results demonstrate that the method achieves substantial efficiency gains while maintaining strong reasoning performance, providing a practical solution for efficient deployment of large language models in mathematical reasoning tasks.
title Efficient Mathematical Reasoning Models via Dynamic Pruning and Knowledge Distillation
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2511.17577