Structured Pruning for Diverse Best-of-N Reasoning Optimization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Nguyen, Hieu Trung, Nguyen, Bao, Nguyen, Viet Anh
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912418717761536
author Nguyen, Hieu Trung
Nguyen, Bao
Nguyen, Viet Anh
author_facet Nguyen, Hieu Trung
Nguyen, Bao
Nguyen, Viet Anh
contents Model pruning in transformer-based language models, traditionally viewed as a means of achieving computational savings, can enhance the model's reasoning capabilities. In this work, we uncover a surprising phenomenon: the selective pruning of certain attention heads leads to improvements in reasoning performance, particularly on challenging tasks. Motivated by this observation, we propose SPRINT, a novel contrastive learning framework that dynamically selects the optimal head and layer to prune during inference. By aligning question embeddings with head embeddings, SPRINT identifies those pruned-head configurations that result in more accurate reasoning. Extensive experiments demonstrate that our method significantly outperforms traditional best-of-$N$ and random head selection strategies on the MATH500 and GSM8K datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2506_03978
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Structured Pruning for Diverse Best-of-N Reasoning Optimization
Nguyen, Hieu Trung
Nguyen, Bao
Nguyen, Viet Anh
Computation and Language
Machine Learning
Model pruning in transformer-based language models, traditionally viewed as a means of achieving computational savings, can enhance the model's reasoning capabilities. In this work, we uncover a surprising phenomenon: the selective pruning of certain attention heads leads to improvements in reasoning performance, particularly on challenging tasks. Motivated by this observation, we propose SPRINT, a novel contrastive learning framework that dynamically selects the optimal head and layer to prune during inference. By aligning question embeddings with head embeddings, SPRINT identifies those pruned-head configurations that result in more accurate reasoning. Extensive experiments demonstrate that our method significantly outperforms traditional best-of-$N$ and random head selection strategies on the MATH500 and GSM8K datasets.
title Structured Pruning for Diverse Best-of-N Reasoning Optimization
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2506.03978