Arena Learning: Build Data Flywheel for LLMs Post-training via Simulated Chatbot Arena

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Haipeng, Sun, Qingfeng, Xu, Can, Zhao, Pu, Lin, Qingwei, Lou, Jianguang, Chen, Shifeng, Tang, Yansong, Chen, Weizhu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913431385276416
author Luo, Haipeng
Sun, Qingfeng
Xu, Can
Zhao, Pu
Lin, Qingwei
Lou, Jianguang
Chen, Shifeng
Tang, Yansong
Chen, Weizhu
author_facet Luo, Haipeng
Sun, Qingfeng
Xu, Can
Zhao, Pu
Lin, Qingwei
Lou, Jianguang
Chen, Shifeng
Tang, Yansong
Chen, Weizhu
contents Assessing the effectiveness of large language models (LLMs) presents substantial challenges. The method of conducting human-annotated battles in an online Chatbot Arena is a highly effective evaluative technique. However, this approach is limited by the costs and time required for human annotation. In this paper, we introduce Arena Learning, an innovative offline strategy designed to simulate these arena battles using AI-driven annotations to evaluate battle outcomes, thus facilitating the continuous improvement of the target model through both supervised fine-tuning and reinforcement learning. Arena Learning comprises two key elements. First, it ensures precise evaluations and maintains consistency between offline simulations and online competitions via WizardArena, a pipeline developed to accurately predict the Elo rankings of various models using a meticulously designed offline test set. Our results demonstrate that WizardArena's predictions closely align with those from the online Arena. Second, it involves the continuous improvement of training data based on the battle results and the refined model. We establish a data flywheel to iteratively update the training data by highlighting the weaknesses of the target model based on its battle results, enabling it to learn from the strengths of multiple different models. We apply Arena Learning to train our target model, WizardLM-$β$, and demonstrate significant performance enhancements across various metrics. This fully automated training and evaluation pipeline sets the stage for continuous advancements in various LLMs via post-training. Notably, Arena Learning plays a pivotal role in the success of WizardLM-2, and this paper serves both as an exploration of its efficacy and a foundational study for future discussions related to WizardLM-2 and its derivatives.
format Preprint
id arxiv_https___arxiv_org_abs_2407_10627
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Arena Learning: Build Data Flywheel for LLMs Post-training via Simulated Chatbot Arena
Luo, Haipeng
Sun, Qingfeng
Xu, Can
Zhao, Pu
Lin, Qingwei
Lou, Jianguang
Chen, Shifeng
Tang, Yansong
Chen, Weizhu
Computation and Language
Artificial Intelligence
Machine Learning
Assessing the effectiveness of large language models (LLMs) presents substantial challenges. The method of conducting human-annotated battles in an online Chatbot Arena is a highly effective evaluative technique. However, this approach is limited by the costs and time required for human annotation. In this paper, we introduce Arena Learning, an innovative offline strategy designed to simulate these arena battles using AI-driven annotations to evaluate battle outcomes, thus facilitating the continuous improvement of the target model through both supervised fine-tuning and reinforcement learning. Arena Learning comprises two key elements. First, it ensures precise evaluations and maintains consistency between offline simulations and online competitions via WizardArena, a pipeline developed to accurately predict the Elo rankings of various models using a meticulously designed offline test set. Our results demonstrate that WizardArena's predictions closely align with those from the online Arena. Second, it involves the continuous improvement of training data based on the battle results and the refined model. We establish a data flywheel to iteratively update the training data by highlighting the weaknesses of the target model based on its battle results, enabling it to learn from the strengths of multiple different models. We apply Arena Learning to train our target model, WizardLM-$β$, and demonstrate significant performance enhancements across various metrics. This fully automated training and evaluation pipeline sets the stage for continuous advancements in various LLMs via post-training. Notably, Arena Learning plays a pivotal role in the success of WizardLM-2, and this paper serves both as an exploration of its efficacy and a foundational study for future discussions related to WizardLM-2 and its derivatives.
title Arena Learning: Build Data Flywheel for LLMs Post-training via Simulated Chatbot Arena
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2407.10627