ARES: Automated Rubric Synthesis for Scalable LLM Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Xiaoyuan, Bao, Keqin, Li, Moxin, Ma, Yubo, Zhang, Yichang, Wang, Wenjie, Feng, Fuli, Liu, Dayiheng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911715211345920
author Li, Xiaoyuan
Bao, Keqin
Li, Moxin
Ma, Yubo
Zhang, Yichang
Wang, Wenjie
Feng, Fuli
Liu, Dayiheng
author_facet Li, Xiaoyuan
Bao, Keqin
Li, Moxin
Ma, Yubo
Zhang, Yichang
Wang, Wenjie
Feng, Fuli
Liu, Dayiheng
contents Rubric-based rewards offer a promising way to extend reinforcement learning (RL) for large language models beyond tasks with automatically verifiable answers. However, scaling rubric-based RL remains challenging: existing approaches often rely on expert-written rubrics and manually constructed question sets, while fixed task-level rubrics may fail to capture the evaluation requirements of individual questions. We propose ARES (Automated Rubric synthEsis for Scalable RL), a framework for automatically constructing rubric-based RL data at scale. Starting from raw pretraining documents, ARES converts source knowledge into self-contained question-answer pairs and co-generates question-specific weighted rubrics, enabling instance-level reward supervision for open-ended responses. To improve diversity and quality, ARES conditions generation on domain labels and persona information, and applies validation filters for question self-containment, answer faithfulness, and rubric validity. Using ARES, we construct 100K rubric-annotated instances across ten domains. Experiments on seven benchmarks show that rubric-based RL trained with ARES, outperforms continual pretraining, supervised fine-tuning, and binary-reward RL, with the largest gains on multi-dimensional open-ended tasks such as healthcare and instruction following.
format Preprint
id arxiv_https___arxiv_org_abs_2605_23454
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ARES: Automated Rubric Synthesis for Scalable LLM Reinforcement Learning
Li, Xiaoyuan
Bao, Keqin
Li, Moxin
Ma, Yubo
Zhang, Yichang
Wang, Wenjie
Feng, Fuli
Liu, Dayiheng
Computation and Language
Rubric-based rewards offer a promising way to extend reinforcement learning (RL) for large language models beyond tasks with automatically verifiable answers. However, scaling rubric-based RL remains challenging: existing approaches often rely on expert-written rubrics and manually constructed question sets, while fixed task-level rubrics may fail to capture the evaluation requirements of individual questions. We propose ARES (Automated Rubric synthEsis for Scalable RL), a framework for automatically constructing rubric-based RL data at scale. Starting from raw pretraining documents, ARES converts source knowledge into self-contained question-answer pairs and co-generates question-specific weighted rubrics, enabling instance-level reward supervision for open-ended responses. To improve diversity and quality, ARES conditions generation on domain labels and persona information, and applies validation filters for question self-containment, answer faithfulness, and rubric validity. Using ARES, we construct 100K rubric-annotated instances across ten domains. Experiments on seven benchmarks show that rubric-based RL trained with ARES, outperforms continual pretraining, supervised fine-tuning, and binary-reward RL, with the largest gains on multi-dimensional open-ended tasks such as healthcare and instruction following.
title ARES: Automated Rubric Synthesis for Scalable LLM Reinforcement Learning
topic Computation and Language
url https://arxiv.org/abs/2605.23454