SibylSense: Adaptive Rubric Learning via Memory Tuning and Adversarial Probing
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915813462638592 |
|---|---|
| author | Xu, Yifei Potje, Guilherme Shandilya, Shivam Yuan, Tiancheng Nunes, Leonardo de Oliveira Agarwal, Rakshanda Asgari, Saeid Atkinson, Adam Kıcıman, Emre Lu, Songwu Chandra, Ranveer Chakraborty, Tusher |
| author_facet | Xu, Yifei Potje, Guilherme Shandilya, Shivam Yuan, Tiancheng Nunes, Leonardo de Oliveira Agarwal, Rakshanda Asgari, Saeid Atkinson, Adam Kıcıman, Emre Lu, Songwu Chandra, Ranveer Chakraborty, Tusher |
| contents | Designing aligned and robust rewards for open-ended generation remains a key barrier to RL post-training. Rubrics provide structured, interpretable supervision, but scaling rubric construction is difficult: expert rubrics are costly, prompted rubrics are often superficial or inconsistent, and fixed-pool discriminative rubrics can saturate and drift, enabling reward hacking. We present SibylSense, an inference-time learning approach that adapts a frozen rubric generator through a tunable memory bank of validated rubric items. Memory is updated via verifier-based item rewards measured by reference-candidate answer discriminative gaps from a handful of examples. SibylSense alternates memory tuning with a rubric-adversarial policy update that produces rubric-satisfying candidate answers, shrinking discriminative gaps and driving the rubric generator to capture new quality dimensions. Experiments on two open-ended tasks show that SibylSense yields more discriminative rubrics and improves downstream RL performance over static and non-adaptive baselines. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_20751 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | SibylSense: Adaptive Rubric Learning via Memory Tuning and Adversarial Probing Xu, Yifei Potje, Guilherme Shandilya, Shivam Yuan, Tiancheng Nunes, Leonardo de Oliveira Agarwal, Rakshanda Asgari, Saeid Atkinson, Adam Kıcıman, Emre Lu, Songwu Chandra, Ranveer Chakraborty, Tusher Computation and Language Artificial Intelligence Machine Learning Designing aligned and robust rewards for open-ended generation remains a key barrier to RL post-training. Rubrics provide structured, interpretable supervision, but scaling rubric construction is difficult: expert rubrics are costly, prompted rubrics are often superficial or inconsistent, and fixed-pool discriminative rubrics can saturate and drift, enabling reward hacking. We present SibylSense, an inference-time learning approach that adapts a frozen rubric generator through a tunable memory bank of validated rubric items. Memory is updated via verifier-based item rewards measured by reference-candidate answer discriminative gaps from a handful of examples. SibylSense alternates memory tuning with a rubric-adversarial policy update that produces rubric-satisfying candidate answers, shrinking discriminative gaps and driving the rubric generator to capture new quality dimensions. Experiments on two open-ended tasks show that SibylSense yields more discriminative rubrics and improves downstream RL performance over static and non-adaptive baselines. |
| title | SibylSense: Adaptive Rubric Learning via Memory Tuning and Adversarial Probing |
| topic | Computation and Language Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2602.20751 |