AMARIS: A Memory-Augmented Rubric Improvement System for Rubric-Based Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Peilin, Zhang, Xinlu, Wan, Kun, Zhao, Wentian, Wu, Gang, Du, Xinya, Chen, Zhiyu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917535907053568
author Wu, Peilin
Zhang, Xinlu
Wan, Kun
Zhao, Wentian
Wu, Gang
Du, Xinya
Chen, Zhiyu
author_facet Wu, Peilin
Zhang, Xinlu
Wan, Kun
Zhao, Wentian
Wu, Gang
Du, Xinya
Chen, Zhiyu
contents Rubric-based reward shaping provides interpretable and editable reward signals for fine-tuning LLMs via reinforcement learning (RL), but existing adaptive rubric methods typically update criteria from local evidence such as the current batch or instance-level comparisons. This local view discards diagnostic information produced during training, making it difficult to track recurring failures, evaluate previous rubric edits, or raise standards once earlier criteria become saturated. We introduce AMARIS, A Memory-Augmented Rubric Improvement System that grounds rubric updates in longitudinal training evidence. AMARIS stores rollout analyses, step-level summaries, and rubric update records in a persistent evaluation memory, then retrieves recent and semantically relevant history to revise rubrics. We evaluate AMARIS across science, medicine, instruction following, and creative writing under both global and instance-specific rubric settings. AMARIS improves over static, local-adaptive, and memory-ablated baselines, such as +2.8 points on GPQA-Diamond and +2.2 points on IFBench over the strongest baselines, while analysis shows that memory reduces oscillatory rubric edits and supports a progression from early failure correction to later curriculum advancement. AMARIS runs asynchronously alongside the normal RL loop, reducing blocking latency relative to synchronous rubric updates.
format Preprint
id arxiv_https___arxiv_org_abs_2605_18592
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AMARIS: A Memory-Augmented Rubric Improvement System for Rubric-Based Reinforcement Learning
Wu, Peilin
Zhang, Xinlu
Wan, Kun
Zhao, Wentian
Wu, Gang
Du, Xinya
Chen, Zhiyu
Machine Learning
Artificial Intelligence
Computation and Language
Rubric-based reward shaping provides interpretable and editable reward signals for fine-tuning LLMs via reinforcement learning (RL), but existing adaptive rubric methods typically update criteria from local evidence such as the current batch or instance-level comparisons. This local view discards diagnostic information produced during training, making it difficult to track recurring failures, evaluate previous rubric edits, or raise standards once earlier criteria become saturated. We introduce AMARIS, A Memory-Augmented Rubric Improvement System that grounds rubric updates in longitudinal training evidence. AMARIS stores rollout analyses, step-level summaries, and rubric update records in a persistent evaluation memory, then retrieves recent and semantically relevant history to revise rubrics. We evaluate AMARIS across science, medicine, instruction following, and creative writing under both global and instance-specific rubric settings. AMARIS improves over static, local-adaptive, and memory-ablated baselines, such as +2.8 points on GPQA-Diamond and +2.2 points on IFBench over the strongest baselines, while analysis shows that memory reduces oscillatory rubric edits and supports a progression from early failure correction to later curriculum advancement. AMARIS runs asynchronously alongside the normal RL loop, reducing blocking latency relative to synchronous rubric updates.
title AMARIS: A Memory-Augmented Rubric Improvement System for Rubric-Based Reinforcement Learning
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2605.18592