Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Maiti, Shalini, Budhiraja, Amar, Gauri, Bhavul, Chaurasia, Gaurav, Protopopov, Anton, Audran-Reiss, Alexis, Slater, Michael, Magka, Despoina, Shavrina, Tatiana, Raileanu, Roberta, Bachrach, Yoram
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909908331397120
author Maiti, Shalini
Budhiraja, Amar
Gauri, Bhavul
Chaurasia, Gaurav
Protopopov, Anton
Audran-Reiss, Alexis
Slater, Michael
Magka, Despoina
Shavrina, Tatiana
Raileanu, Roberta
Bachrach, Yoram
author_facet Maiti, Shalini
Budhiraja, Amar
Gauri, Bhavul
Chaurasia, Gaurav
Protopopov, Anton
Audran-Reiss, Alexis
Slater, Michael
Magka, Despoina
Shavrina, Tatiana
Raileanu, Roberta
Bachrach, Yoram
contents Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse domains, but their training remains resource- and time-intensive, requiring massive compute power and careful orchestration of training procedures. Model souping-the practice of averaging weights from multiple models of the same architecture-has emerged as a promising pre- and post-training technique that can enhance performance without expensive retraining. In this paper, we introduce Soup Of Category Experts (SoCE), a principled approach for model souping that utilizes benchmark composition to identify optimal model candidates and applies non-uniform weighted averaging to maximize performance. Contrary to previous uniform-averaging approaches, our method leverages the observation that benchmark categories often exhibit low inter-correlations in model performance. SoCE identifies "expert" models for each weakly-correlated category cluster and combines them using optimized weighted averaging rather than uniform weights. We demonstrate that the proposed method improves performance and robustness across multiple domains, including multilingual capabilities, tool calling, and math and achieves state-of-the-art results on the Berkeley Function Calling Leaderboard.
format Preprint
id arxiv_https___arxiv_org_abs_2511_13254
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance
Maiti, Shalini
Budhiraja, Amar
Gauri, Bhavul
Chaurasia, Gaurav
Protopopov, Anton
Audran-Reiss, Alexis
Slater, Michael
Magka, Despoina
Shavrina, Tatiana
Raileanu, Roberta
Bachrach, Yoram
Computation and Language
Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse domains, but their training remains resource- and time-intensive, requiring massive compute power and careful orchestration of training procedures. Model souping-the practice of averaging weights from multiple models of the same architecture-has emerged as a promising pre- and post-training technique that can enhance performance without expensive retraining. In this paper, we introduce Soup Of Category Experts (SoCE), a principled approach for model souping that utilizes benchmark composition to identify optimal model candidates and applies non-uniform weighted averaging to maximize performance. Contrary to previous uniform-averaging approaches, our method leverages the observation that benchmark categories often exhibit low inter-correlations in model performance. SoCE identifies "expert" models for each weakly-correlated category cluster and combines them using optimized weighted averaging rather than uniform weights. We demonstrate that the proposed method improves performance and robustness across multiple domains, including multilingual capabilities, tool calling, and math and achieves state-of-the-art results on the Berkeley Function Calling Leaderboard.
title Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance
topic Computation and Language
url https://arxiv.org/abs/2511.13254