$n$-Musketeers: Reinforcement Learning Shapes Collaboration Among Language Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Masukawa, Ryozo, Yun, Sanggeon, Oh, Hyunwoo, Jeong, SuhgHeon, Hassa, Raheeb, Chen, Hanning, Huang, Wenjun, Imani, Mahdi, Mercati, Pietro, Bastian, Nathaniel D., Imani, Mohsen
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918330226442240
author Masukawa, Ryozo
Yun, Sanggeon
Oh, Hyunwoo
Jeong, SuhgHeon
Hassa, Raheeb
Chen, Hanning
Huang, Wenjun
Imani, Mahdi
Mercati, Pietro
Bastian, Nathaniel D.
Imani, Mohsen
author_facet Masukawa, Ryozo
Yun, Sanggeon
Oh, Hyunwoo
Jeong, SuhgHeon
Hassa, Raheeb
Chen, Hanning
Huang, Wenjun
Imani, Mahdi
Mercati, Pietro
Bastian, Nathaniel D.
Imani, Mohsen
contents Recent progress in reinforcement learning with verifiable rewards (RLVR) shows that small, specialized language models (SLMs) can exhibit structured reasoning without relying on large monolithic LLMs. We introduce soft hidden-state collaboration, where multiple heterogeneous frozen SLM experts are integrated through their internal representations via a trainable attention interface. Experiments on Reasoning Gym and GSM8K show that this latent integration is competitive with strong single-model RLVR baselines. Ablations further reveal a dual mechanism of expert utilization: for simpler arithmetic domains, performance gains can largely be explained by static expert preferences, whereas more challenging settings induce increasingly concentrated and structured expert attention over training, indicating emergent specialization in how the router connects to relevant experts. Overall, hidden-state collaboration provides a compact mechanism for leveraging frozen experts, while offering an observational window into expert utilization patterns and their evolution under RLVR.
format Preprint
id arxiv_https___arxiv_org_abs_2602_09173
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle $n$-Musketeers: Reinforcement Learning Shapes Collaboration Among Language Models
Masukawa, Ryozo
Yun, Sanggeon
Oh, Hyunwoo
Jeong, SuhgHeon
Hassa, Raheeb
Chen, Hanning
Huang, Wenjun
Imani, Mahdi
Mercati, Pietro
Bastian, Nathaniel D.
Imani, Mohsen
Machine Learning
Artificial Intelligence
Recent progress in reinforcement learning with verifiable rewards (RLVR) shows that small, specialized language models (SLMs) can exhibit structured reasoning without relying on large monolithic LLMs. We introduce soft hidden-state collaboration, where multiple heterogeneous frozen SLM experts are integrated through their internal representations via a trainable attention interface. Experiments on Reasoning Gym and GSM8K show that this latent integration is competitive with strong single-model RLVR baselines. Ablations further reveal a dual mechanism of expert utilization: for simpler arithmetic domains, performance gains can largely be explained by static expert preferences, whereas more challenging settings induce increasingly concentrated and structured expert attention over training, indicating emergent specialization in how the router connects to relevant experts. Overall, hidden-state collaboration provides a compact mechanism for leveraging frozen experts, while offering an observational window into expert utilization patterns and their evolution under RLVR.
title $n$-Musketeers: Reinforcement Learning Shapes Collaboration Among Language Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.09173