Weight Ensembling Improves Reasoning in Language Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Dang, Xingyu, Baek, Christina, Wen, Kaiyue, Kolter, Zico, Raghunathan, Aditi
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908580016291840
author Dang, Xingyu
Baek, Christina
Wen, Kaiyue
Kolter, Zico
Raghunathan, Aditi
author_facet Dang, Xingyu
Baek, Christina
Wen, Kaiyue
Kolter, Zico
Raghunathan, Aditi
contents We investigate a failure mode that arises during the training of reasoning models, where the diversity of generations begins to collapse, leading to suboptimal test-time scaling. Notably, the Pass@1 rate reliably improves during supervised finetuning (SFT), but Pass@k rapidly deteriorates. Surprisingly, a simple intervention of interpolating the weights of the latest SFT checkpoint with an early checkpoint, otherwise known as WiSE-FT, almost completely recovers Pass@k while also improving Pass@1. The WiSE-FT variant achieves better test-time scaling (Best@k, majority vote) and achieves superior results with less data when tuned further by reinforcement learning. Finally, we find that WiSE-FT provides complementary performance gains that cannot be achieved only through diversity-inducing decoding strategies, like temperature scaling. We formalize a bias-variance tradeoff of Pass@k with respect to the expectation and variance of Pass@1 over the test distribution. We find that WiSE-FT can reduce bias and variance simultaneously, while temperature scaling inherently trades off between bias and variance.
format Preprint
id arxiv_https___arxiv_org_abs_2504_10478
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Weight Ensembling Improves Reasoning in Language Models
Dang, Xingyu
Baek, Christina
Wen, Kaiyue
Kolter, Zico
Raghunathan, Aditi
Machine Learning
Artificial Intelligence
We investigate a failure mode that arises during the training of reasoning models, where the diversity of generations begins to collapse, leading to suboptimal test-time scaling. Notably, the Pass@1 rate reliably improves during supervised finetuning (SFT), but Pass@k rapidly deteriorates. Surprisingly, a simple intervention of interpolating the weights of the latest SFT checkpoint with an early checkpoint, otherwise known as WiSE-FT, almost completely recovers Pass@k while also improving Pass@1. The WiSE-FT variant achieves better test-time scaling (Best@k, majority vote) and achieves superior results with less data when tuned further by reinforcement learning. Finally, we find that WiSE-FT provides complementary performance gains that cannot be achieved only through diversity-inducing decoding strategies, like temperature scaling. We formalize a bias-variance tradeoff of Pass@k with respect to the expectation and variance of Pass@1 over the test distribution. We find that WiSE-FT can reduce bias and variance simultaneously, while temperature scaling inherently trades off between bias and variance.
title Weight Ensembling Improves Reasoning in Language Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2504.10478