RLAE: Reinforcement Learning-Assisted Ensemble for LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fu, Yuqian, Zhu, Yuanheng, Chai, Jiajun, Yin, Guojun, Lin, Wei, Zhang, Qichao, Zhao, Dongbin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910978085486592
author Fu, Yuqian
Zhu, Yuanheng
Chai, Jiajun
Yin, Guojun
Lin, Wei
Zhang, Qichao
Zhao, Dongbin
author_facet Fu, Yuqian
Zhu, Yuanheng
Chai, Jiajun
Yin, Guojun
Lin, Wei
Zhang, Qichao
Zhao, Dongbin
contents Ensembling large language models (LLMs) can effectively combine diverse strengths of different models, offering a promising approach to enhance performance across various tasks. However, existing methods typically rely on fixed weighting strategies that fail to adapt to the dynamic, context-dependent characteristics of LLM capabilities. In this work, we propose Reinforcement Learning-Assisted Ensemble for LLMs (RLAE), a novel framework that reformulates LLM ensemble through the lens of a Markov Decision Process (MDP). Our approach introduces a RL agent that dynamically adjusts ensemble weights by considering both input context and intermediate generation states, with the agent being trained using rewards that directly correspond to the quality of final outputs. We implement RLAE using both single-agent and multi-agent reinforcement learning algorithms ($\text{RLAE}_\text{PPO}$ and $\text{RLAE}_\text{MAPPO}$ ), demonstrating substantial improvements over conventional ensemble methods. Extensive evaluations on a diverse set of tasks show that RLAE outperforms existing approaches by up to $3.3\%$ accuracy points, offering a more effective framework for LLM ensembling. Furthermore, our method exhibits superior generalization capabilities across different tasks without the need for retraining, while simultaneously achieving lower time latency.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00439
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RLAE: Reinforcement Learning-Assisted Ensemble for LLMs
Fu, Yuqian
Zhu, Yuanheng
Chai, Jiajun
Yin, Guojun
Lin, Wei
Zhang, Qichao
Zhao, Dongbin
Machine Learning
Artificial Intelligence
Ensembling large language models (LLMs) can effectively combine diverse strengths of different models, offering a promising approach to enhance performance across various tasks. However, existing methods typically rely on fixed weighting strategies that fail to adapt to the dynamic, context-dependent characteristics of LLM capabilities. In this work, we propose Reinforcement Learning-Assisted Ensemble for LLMs (RLAE), a novel framework that reformulates LLM ensemble through the lens of a Markov Decision Process (MDP). Our approach introduces a RL agent that dynamically adjusts ensemble weights by considering both input context and intermediate generation states, with the agent being trained using rewards that directly correspond to the quality of final outputs. We implement RLAE using both single-agent and multi-agent reinforcement learning algorithms ($\text{RLAE}_\text{PPO}$ and $\text{RLAE}_\text{MAPPO}$ ), demonstrating substantial improvements over conventional ensemble methods. Extensive evaluations on a diverse set of tasks show that RLAE outperforms existing approaches by up to $3.3\%$ accuracy points, offering a more effective framework for LLM ensembling. Furthermore, our method exhibits superior generalization capabilities across different tasks without the need for retraining, while simultaneously achieving lower time latency.
title RLAE: Reinforcement Learning-Assisted Ensemble for LLMs
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.00439