Arena Learning: Build Data Flywheel for LLMs Post-training via Simulated Chatbot Arena
Fuente:
arXiv
Saved in:
| Main Authors: | Luo, Haipeng, Sun, Qingfeng, Xu, Can, Zhao, Pu, Lin, Qingwei, Lou, Jianguang, Chen, Shifeng, Tang, Yansong, Chen, Weizhu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct
by: Luo, Haipeng, et al.
Published: (2023)
by: Luo, Haipeng, et al.
Published: (2023)
AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task Generation
by: Hu, Mengkang, et al.
Published: (2024)
by: Hu, Mengkang, et al.
Published: (2024)
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
by: Chiang, Wei-Lin, et al.
Published: (2024)
by: Chiang, Wei-Lin, et al.
Published: (2024)
OpenDataArena: A Fair and Open Arena for Benchmarking Post-Training Dataset Value
by: Cai, Mengzhang, et al.
Published: (2025)
by: Cai, Mengzhang, et al.
Published: (2025)
CHARM: Calibrating Reward Models With Chatbot Arena Scores
by: Zhu, Xiao, et al.
Published: (2025)
by: Zhu, Xiao, et al.
Published: (2025)
Improving Your Model Ranking on Chatbot Arena by Vote Rigging
by: Min, Rui, et al.
Published: (2025)
by: Min, Rui, et al.
Published: (2025)
WizardLM: Empowering large pre-trained language models to follow complex instructions
by: Xu, Can, et al.
Published: (2023)
by: Xu, Can, et al.
Published: (2023)
Chatbot Arena Meets Nuggets: Towards Explanations and Diagnostics in the Evaluation of LLM Responses
by: Sharifymoghaddam, Sahel, et al.
Published: (2025)
by: Sharifymoghaddam, Sahel, et al.
Published: (2025)
Will LLMs be Professional at Fund Investment? DeepFund: A Live Arena Perspective
by: Li, Changlun, et al.
Published: (2025)
by: Li, Changlun, et al.
Published: (2025)
VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation
by: Luo, Ziyang, et al.
Published: (2024)
by: Luo, Ziyang, et al.
Published: (2024)
BuildArena: A Physics-Aligned Interactive Benchmark of LLMs for Engineering Construction
by: Xia, Tian, et al.
Published: (2025)
by: Xia, Tian, et al.
Published: (2025)
Search Arena: Analyzing Search-Augmented LLMs
by: Miroyan, Mihran, et al.
Published: (2025)
by: Miroyan, Mihran, et al.
Published: (2025)
Arenas
Published: (2010)
Published: (2010)
AgentMath: Empowering Mathematical Reasoning for Large Language Models via Tool-Augmented Agent
by: Luo, Haipeng, et al.
Published: (2025)
by: Luo, Haipeng, et al.
Published: (2025)
BABE: Biology Arena BEnchmark
by: Zhou, Junting, et al.
Published: (2026)
by: Zhou, Junting, et al.
Published: (2026)
Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
by: Dekoninck, Jasper, et al.
Published: (2026)
by: Dekoninck, Jasper, et al.
Published: (2026)
LLM Content Moderation and User Satisfaction: Evidence from Response Refusals in Chatbot Arena
by: Pasch, Stefan
Published: (2025)
by: Pasch, Stefan
Published: (2025)
Players and Arenas
Published: (2025)
Published: (2025)
Arena Hukum
Published: (2023)
Published: (2023)
Linguarum Arena
Published: (2011)
Published: (2011)
Players and Arenas
Published: (2019)
Published: (2019)
TextArena
by: Guertler, Leon, et al.
Published: (2025)
by: Guertler, Leon, et al.
Published: (2025)
Arena sagrada
by: Karen E. Lange
Published: (2007)
by: Karen E. Lange
Published: (2007)
SwingArena: Competitive Programming Arena for Long-context GitHub Issue Solving
by: Xu, Wendong, et al.
Published: (2025)
by: Xu, Wendong, et al.
Published: (2025)
World Reasoning Arena
by: PAN Team, et al.
Published: (2026)
by: PAN Team, et al.
Published: (2026)
Automatic Instruction Evolving for Large Language Models
by: Zeng, Weihao, et al.
Published: (2024)
by: Zeng, Weihao, et al.
Published: (2024)
WizardCoder: Empowering Code Large Language Models with Evol-Instruct
by: Luo, Ziyang, et al.
Published: (2023)
by: Luo, Ziyang, et al.
Published: (2023)
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
by: Balunović, Mislav, et al.
Published: (2025)
by: Balunović, Mislav, et al.
Published: (2025)
Towards a Unified Paradigm: Integrating Recommendation Systems as a New Language in Large Models
by: Zheng, Kai, et al.
Published: (2024)
by: Zheng, Kai, et al.
Published: (2024)
HomeFlow: A Data Flywheel for Smart Home Agent Training with Verifiable Simulation
by: Gu, Yi, et al.
Published: (2026)
by: Gu, Yi, et al.
Published: (2026)
CharacterFlywheel: Scaling Iterative Improvement of Engaging and Steerable LLMs in Production
by: Nie, Yixin, et al.
Published: (2026)
by: Nie, Yixin, et al.
Published: (2026)
WebArena: A Realistic Web Environment for Building Autonomous Agents
by: Zhou, Shuyan, et al.
Published: (2023)
by: Zhou, Shuyan, et al.
Published: (2023)
Israel’s Academic Arena
by: Haliwa, Pinhas
Published: (2021)
by: Haliwa, Pinhas
Published: (2021)
Cataloguing in the International Arena.
by: Cook, C. Donald
Published: (1986)
by: Cook, C. Donald
Published: (1986)
GraphArena: Evaluating and Exploring Large Language Models on Graph Computation
by: Tang, Jianheng, et al.
Published: (2024)
by: Tang, Jianheng, et al.
Published: (2024)
MedArena: Comparing LLMs for Medicine-in-the-Wild Clinician Preferences
by: Wu, Eric, et al.
Published: (2026)
by: Wu, Eric, et al.
Published: (2026)
CalArena: A Large-Scale Post-Hoc Calibration Benchmark
by: Berta, Eugène, et al.
Published: (2026)
by: Berta, Eugène, et al.
Published: (2026)
TimeArena: Shaping Efficient Multitasking Language Agents in a Time-Aware Simulation
by: Zhang, Yikai, et al.
Published: (2024)
by: Zhang, Yikai, et al.
Published: (2024)
VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks
by: Jang, Lawrence, et al.
Published: (2024)
by: Jang, Lawrence, et al.
Published: (2024)
GenAI Arena: An Open Evaluation Platform for Generative Models
by: Jiang, Dongfu, et al.
Published: (2024)
by: Jiang, Dongfu, et al.
Published: (2024)
Similar Items
-
WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct
by: Luo, Haipeng, et al.
Published: (2023) -
AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task Generation
by: Hu, Mengkang, et al.
Published: (2024) -
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
by: Chiang, Wei-Lin, et al.
Published: (2024) -
OpenDataArena: A Fair and Open Arena for Benchmarking Post-Training Dataset Value
by: Cai, Mengzhang, et al.
Published: (2025) -
CHARM: Calibrating Reward Models With Chatbot Arena Scores
by: Zhu, Xiao, et al.
Published: (2025)