Structured Pruning for Diverse Best-of-N Reasoning Optimization
Fuente:
arXiv
Salvato in:
| Autori principali: | , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866912418717761536 |
|---|---|
| author | Nguyen, Hieu Trung Nguyen, Bao Nguyen, Viet Anh |
| author_facet | Nguyen, Hieu Trung Nguyen, Bao Nguyen, Viet Anh |
| contents | Model pruning in transformer-based language models, traditionally viewed as a means of achieving computational savings, can enhance the model's reasoning capabilities. In this work, we uncover a surprising phenomenon: the selective pruning of certain attention heads leads to improvements in reasoning performance, particularly on challenging tasks. Motivated by this observation, we propose SPRINT, a novel contrastive learning framework that dynamically selects the optimal head and layer to prune during inference. By aligning question embeddings with head embeddings, SPRINT identifies those pruned-head configurations that result in more accurate reasoning. Extensive experiments demonstrate that our method significantly outperforms traditional best-of-$N$ and random head selection strategies on the MATH500 and GSM8K datasets. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_03978 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Structured Pruning for Diverse Best-of-N Reasoning Optimization Nguyen, Hieu Trung Nguyen, Bao Nguyen, Viet Anh Computation and Language Machine Learning Model pruning in transformer-based language models, traditionally viewed as a means of achieving computational savings, can enhance the model's reasoning capabilities. In this work, we uncover a surprising phenomenon: the selective pruning of certain attention heads leads to improvements in reasoning performance, particularly on challenging tasks. Motivated by this observation, we propose SPRINT, a novel contrastive learning framework that dynamically selects the optimal head and layer to prune during inference. By aligning question embeddings with head embeddings, SPRINT identifies those pruned-head configurations that result in more accurate reasoning. Extensive experiments demonstrate that our method significantly outperforms traditional best-of-$N$ and random head selection strategies on the MATH500 and GSM8K datasets. |
| title | Structured Pruning for Diverse Best-of-N Reasoning Optimization |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2506.03978 |