SibylSense: Adaptive Rubric Learning via Memory Tuning and Adversarial Probing

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xu, Yifei, Potje, Guilherme, Shandilya, Shivam, Yuan, Tiancheng, Nunes, Leonardo de Oliveira, Agarwal, Rakshanda, Asgari, Saeid, Atkinson, Adam, Kıcıman, Emre, Lu, Songwu, Chandra, Ranveer, Chakraborty, Tusher
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915813462638592
author Xu, Yifei
Potje, Guilherme
Shandilya, Shivam
Yuan, Tiancheng
Nunes, Leonardo de Oliveira
Agarwal, Rakshanda
Asgari, Saeid
Atkinson, Adam
Kıcıman, Emre
Lu, Songwu
Chandra, Ranveer
Chakraborty, Tusher
author_facet Xu, Yifei
Potje, Guilherme
Shandilya, Shivam
Yuan, Tiancheng
Nunes, Leonardo de Oliveira
Agarwal, Rakshanda
Asgari, Saeid
Atkinson, Adam
Kıcıman, Emre
Lu, Songwu
Chandra, Ranveer
Chakraborty, Tusher
contents Designing aligned and robust rewards for open-ended generation remains a key barrier to RL post-training. Rubrics provide structured, interpretable supervision, but scaling rubric construction is difficult: expert rubrics are costly, prompted rubrics are often superficial or inconsistent, and fixed-pool discriminative rubrics can saturate and drift, enabling reward hacking. We present SibylSense, an inference-time learning approach that adapts a frozen rubric generator through a tunable memory bank of validated rubric items. Memory is updated via verifier-based item rewards measured by reference-candidate answer discriminative gaps from a handful of examples. SibylSense alternates memory tuning with a rubric-adversarial policy update that produces rubric-satisfying candidate answers, shrinking discriminative gaps and driving the rubric generator to capture new quality dimensions. Experiments on two open-ended tasks show that SibylSense yields more discriminative rubrics and improves downstream RL performance over static and non-adaptive baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2602_20751
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SibylSense: Adaptive Rubric Learning via Memory Tuning and Adversarial Probing
Xu, Yifei
Potje, Guilherme
Shandilya, Shivam
Yuan, Tiancheng
Nunes, Leonardo de Oliveira
Agarwal, Rakshanda
Asgari, Saeid
Atkinson, Adam
Kıcıman, Emre
Lu, Songwu
Chandra, Ranveer
Chakraborty, Tusher
Computation and Language
Artificial Intelligence
Machine Learning
Designing aligned and robust rewards for open-ended generation remains a key barrier to RL post-training. Rubrics provide structured, interpretable supervision, but scaling rubric construction is difficult: expert rubrics are costly, prompted rubrics are often superficial or inconsistent, and fixed-pool discriminative rubrics can saturate and drift, enabling reward hacking. We present SibylSense, an inference-time learning approach that adapts a frozen rubric generator through a tunable memory bank of validated rubric items. Memory is updated via verifier-based item rewards measured by reference-candidate answer discriminative gaps from a handful of examples. SibylSense alternates memory tuning with a rubric-adversarial policy update that produces rubric-satisfying candidate answers, shrinking discriminative gaps and driving the rubric generator to capture new quality dimensions. Experiments on two open-ended tasks show that SibylSense yields more discriminative rubrics and improves downstream RL performance over static and non-adaptive baselines.
title SibylSense: Adaptive Rubric Learning via Memory Tuning and Adversarial Probing
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2602.20751