The Avengers: A Simple Recipe for Uniting Smaller Language Models to Challenge Proprietary Giants

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Yiqun, Li, Hao, Wang, Chenxu, Chen, Linyao, Zhang, Qiaosheng, Ye, Peng, Feng, Shi, Wang, Daling, Wang, Zhen, Wang, Xinrun, Xu, Jia, Bai, Lei, Ouyang, Wanli, Hu, Shuyue
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913899733843968
author Zhang, Yiqun
Li, Hao
Wang, Chenxu
Chen, Linyao
Zhang, Qiaosheng
Ye, Peng
Feng, Shi
Wang, Daling
Wang, Zhen
Wang, Xinrun
Xu, Jia
Bai, Lei
Ouyang, Wanli
Hu, Shuyue
author_facet Zhang, Yiqun
Li, Hao
Wang, Chenxu
Chen, Linyao
Zhang, Qiaosheng
Ye, Peng
Feng, Shi
Wang, Daling
Wang, Zhen
Wang, Xinrun
Xu, Jia
Bai, Lei
Ouyang, Wanli
Hu, Shuyue
contents Proprietary giants are increasingly dominating the race for ever-larger language models. Can open-source, smaller models remain competitive across a broad range of tasks? In this paper, we present the Avengers -- a simple recipe that leverages the collective intelligence of these smaller models. The Avengers builds upon four lightweight operations: (i) embedding: encode queries using a text embedding model; (ii) clustering: group queries based on their semantic similarity; (iii) scoring: scores each model's performance within each cluster; and (iv) voting: improve outputs via repeated sampling and voting. At inference time, each query is embedded and assigned to its nearest cluster. The top-performing model(s) within that cluster are selected to generate the response with repeated sampling. Remarkably, with 10 open-source models (~7B parameters each), the Avengers surpasses GPT-4o, 4.1, and 4.5 in average performance across 15 diverse datasets spanning mathematics, coding, logical reasoning, general knowledge, and affective tasks. In particular, it surpasses GPT-4.1 on mathematics tasks by 18.21% and on code tasks by 7.46%. Furthermore, the Avengers delivers superior out-of-distribution generalization, and remains robust across various embedding models, clustering algorithms, ensemble strategies, and values of its sole parameter -- the number of clusters.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19797
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Avengers: A Simple Recipe for Uniting Smaller Language Models to Challenge Proprietary Giants
Zhang, Yiqun
Li, Hao
Wang, Chenxu
Chen, Linyao
Zhang, Qiaosheng
Ye, Peng
Feng, Shi
Wang, Daling
Wang, Zhen
Wang, Xinrun
Xu, Jia
Bai, Lei
Ouyang, Wanli
Hu, Shuyue
Computation and Language
Proprietary giants are increasingly dominating the race for ever-larger language models. Can open-source, smaller models remain competitive across a broad range of tasks? In this paper, we present the Avengers -- a simple recipe that leverages the collective intelligence of these smaller models. The Avengers builds upon four lightweight operations: (i) embedding: encode queries using a text embedding model; (ii) clustering: group queries based on their semantic similarity; (iii) scoring: scores each model's performance within each cluster; and (iv) voting: improve outputs via repeated sampling and voting. At inference time, each query is embedded and assigned to its nearest cluster. The top-performing model(s) within that cluster are selected to generate the response with repeated sampling. Remarkably, with 10 open-source models (~7B parameters each), the Avengers surpasses GPT-4o, 4.1, and 4.5 in average performance across 15 diverse datasets spanning mathematics, coding, logical reasoning, general knowledge, and affective tasks. In particular, it surpasses GPT-4.1 on mathematics tasks by 18.21% and on code tasks by 7.46%. Furthermore, the Avengers delivers superior out-of-distribution generalization, and remains robust across various embedding models, clustering algorithms, ensemble strategies, and values of its sole parameter -- the number of clusters.
title The Avengers: A Simple Recipe for Uniting Smaller Language Models to Challenge Proprietary Giants
topic Computation and Language
url https://arxiv.org/abs/2505.19797