Wisdom of the Crowd: Reinforcement Learning from Coevolutionary Collective Feedback

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yuan, Wenzhen, Tang, Shengji, Lin, Weihao, Ruan, Jiacheng, Cui, Ganqu, Zhang, Bo, Chen, Tao, Liu, Ting, Fu, Yuzhuo, Ye, Peng, Bai, Lei
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918126438842368
author Yuan, Wenzhen
Tang, Shengji
Lin, Weihao
Ruan, Jiacheng
Cui, Ganqu
Zhang, Bo
Chen, Tao
Liu, Ting
Fu, Yuzhuo
Ye, Peng
Bai, Lei
author_facet Yuan, Wenzhen
Tang, Shengji
Lin, Weihao
Ruan, Jiacheng
Cui, Ganqu
Zhang, Bo
Chen, Tao
Liu, Ting
Fu, Yuzhuo
Ye, Peng
Bai, Lei
contents Reinforcement learning (RL) has significantly enhanced the reasoning capabilities of large language models (LLMs), but its reliance on expensive human-labeled data or complex reward models severely limits scalability. While existing self-feedback methods aim to address this problem, they are constrained by the capabilities of a single model, which can lead to overconfidence in incorrect answers, reward hacking, and even training collapse. To this end, we propose Reinforcement Learning from Coevolutionary Collective Feedback (RLCCF), a novel RL framework that enables multi-model collaborative evolution without external supervision. Specifically, RLCCF optimizes the ability of a model collective by maximizing its Collective Consistency (CC), which jointly trains a diverse ensemble of LLMs and provides reward signals by voting on collective outputs. Moreover, each model's vote is weighted by its Self-Consistency (SC) score, ensuring that more confident models contribute more to the collective decision. Benefiting from the diverse output distributions and complementary abilities of multiple LLMs, RLCCF enables the model collective to continuously enhance its reasoning ability through coevolution. Experiments on four mainstream open-source LLMs across four mathematical reasoning benchmarks demonstrate that our framework yields significant performance gains, achieving an average relative improvement of 16.72\% in accuracy. Notably, RLCCF not only improves the performance of individual models but also enhances the group's majority-voting accuracy by 4.51\%, demonstrating its ability to extend the collective capability boundary of the model collective.
format Preprint
id arxiv_https___arxiv_org_abs_2508_12338
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Wisdom of the Crowd: Reinforcement Learning from Coevolutionary Collective Feedback
Yuan, Wenzhen
Tang, Shengji
Lin, Weihao
Ruan, Jiacheng
Cui, Ganqu
Zhang, Bo
Chen, Tao
Liu, Ting
Fu, Yuzhuo
Ye, Peng
Bai, Lei
Artificial Intelligence
Reinforcement learning (RL) has significantly enhanced the reasoning capabilities of large language models (LLMs), but its reliance on expensive human-labeled data or complex reward models severely limits scalability. While existing self-feedback methods aim to address this problem, they are constrained by the capabilities of a single model, which can lead to overconfidence in incorrect answers, reward hacking, and even training collapse. To this end, we propose Reinforcement Learning from Coevolutionary Collective Feedback (RLCCF), a novel RL framework that enables multi-model collaborative evolution without external supervision. Specifically, RLCCF optimizes the ability of a model collective by maximizing its Collective Consistency (CC), which jointly trains a diverse ensemble of LLMs and provides reward signals by voting on collective outputs. Moreover, each model's vote is weighted by its Self-Consistency (SC) score, ensuring that more confident models contribute more to the collective decision. Benefiting from the diverse output distributions and complementary abilities of multiple LLMs, RLCCF enables the model collective to continuously enhance its reasoning ability through coevolution. Experiments on four mainstream open-source LLMs across four mathematical reasoning benchmarks demonstrate that our framework yields significant performance gains, achieving an average relative improvement of 16.72\% in accuracy. Notably, RLCCF not only improves the performance of individual models but also enhances the group's majority-voting accuracy by 4.51\%, demonstrating its ability to extend the collective capability boundary of the model collective.
title Wisdom of the Crowd: Reinforcement Learning from Coevolutionary Collective Feedback
topic Artificial Intelligence
url https://arxiv.org/abs/2508.12338