Rethinking Rubric Generation for Improving LLM Judge and Reward Modeling for Open-ended Tasks
Fuente:
arXiv
Salvato in:
| Autori principali: | Shen, William F., Qiu, Xinchi, Whitehouse, Chenxi, Alazraki, Lisa, Goel, Shashwat, Barbieri, Francesco, Willi, Timon, Mathur, Akhil, Leontiadis, Ilias |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Training AI Co-Scientists Using Rubric Rewards
di: Goel, Shashwat, et al.
Pubblicazione: (2025)
di: Goel, Shashwat, et al.
Pubblicazione: (2025)
Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
di: Collot, Stephane, et al.
Pubblicazione: (2025)
di: Collot, Stephane, et al.
Pubblicazione: (2025)
Evaluating Privacy Leakage in Split Learning
di: Qiu, Xinchi, et al.
Pubblicazione: (2023)
di: Qiu, Xinchi, et al.
Pubblicazione: (2023)
Scaling Small Agents Through Strategy Auctions
di: Alazraki, Lisa, et al.
Pubblicazione: (2026)
di: Alazraki, Lisa, et al.
Pubblicazione: (2026)
Meta-Reasoning Improves Tool Use in Large Language Models
di: Alazraki, Lisa, et al.
Pubblicazione: (2024)
di: Alazraki, Lisa, et al.
Pubblicazione: (2024)
AdaRubric: Task-Adaptive Rubrics for Reliable LLM Agent Evaluation and Reward Learning
di: Ding, Liang
Pubblicazione: (2026)
di: Ding, Liang
Pubblicazione: (2026)
OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM Alignment
di: Liu, Tianci, et al.
Pubblicazione: (2025)
di: Liu, Tianci, et al.
Pubblicazione: (2025)
Culturally Grounded Physical Commonsense Reasoning in Italian and English: A Submission to the MRL 2025 Shared Task
di: De Santis, Marco, et al.
Pubblicazione: (2025)
di: De Santis, Marco, et al.
Pubblicazione: (2025)
Towards Knowledge-Grounded Natural Language Understanding and Generation
di: Whitehouse, Chenxi
Pubblicazione: (2024)
di: Whitehouse, Chenxi
Pubblicazione: (2024)
Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation
di: Radharapu, Bhaktipriya, et al.
Pubblicazione: (2025)
di: Radharapu, Bhaktipriya, et al.
Pubblicazione: (2025)
Rubric-Grounded RL: Structured Judge Rewards for Generalizable Reasoning
di: Bhattarai, Manish, et al.
Pubblicazione: (2026)
di: Bhattarai, Manish, et al.
Pubblicazione: (2026)
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
di: Whitehouse, Chenxi, et al.
Pubblicazione: (2025)
di: Whitehouse, Chenxi, et al.
Pubblicazione: (2025)
Step-wise Rubric Rewards for LLM Reasoning
di: Xie, Weichu, et al.
Pubblicazione: (2026)
di: Xie, Weichu, et al.
Pubblicazione: (2026)
RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following
di: Pan, Tianjun, et al.
Pubblicazione: (2026)
di: Pan, Tianjun, et al.
Pubblicazione: (2026)
Rubrics to Tokens: Bridging Response-level Rubrics and Token-level Rewards in Instruction Following Tasks
di: Xu, Tianze, et al.
Pubblicazione: (2026)
di: Xu, Tianze, et al.
Pubblicazione: (2026)
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
di: Ding, Ruomeng, et al.
Pubblicazione: (2026)
di: Ding, Ruomeng, et al.
Pubblicazione: (2026)
Enhancing LLM Robustness to Perturbed Instructions: An Empirical Study
di: Agrawal, Aryan, et al.
Pubblicazione: (2025)
di: Agrawal, Aryan, et al.
Pubblicazione: (2025)
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind
di: Lupu, Andrei, et al.
Pubblicazione: (2025)
di: Lupu, Andrei, et al.
Pubblicazione: (2025)
Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards
di: Yuan, Youliang, et al.
Pubblicazione: (2025)
di: Yuan, Youliang, et al.
Pubblicazione: (2025)
Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text
di: Huang, Chengyu, et al.
Pubblicazione: (2026)
di: Huang, Chengyu, et al.
Pubblicazione: (2026)
Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning
di: Ye, Zhiling, et al.
Pubblicazione: (2025)
di: Ye, Zhiling, et al.
Pubblicazione: (2025)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
di: Hong, Yihan, et al.
Pubblicazione: (2026)
di: Hong, Yihan, et al.
Pubblicazione: (2026)
Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR
di: Tyagi, Utkarsh, et al.
Pubblicazione: (2026)
di: Tyagi, Utkarsh, et al.
Pubblicazione: (2026)
AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning
di: Jia, Mengzhao, et al.
Pubblicazione: (2025)
di: Jia, Mengzhao, et al.
Pubblicazione: (2025)
A Judge-free LLM Open-ended Generation Benchmark Based on the Distributional Hypothesis
di: Imajo, Kentaro, et al.
Pubblicazione: (2025)
di: Imajo, Kentaro, et al.
Pubblicazione: (2025)
ProBench: Judging Multimodal Foundation Models on Open-ended Multi-domain Expert Tasks
di: Yang, Yan, et al.
Pubblicazione: (2025)
di: Yang, Yan, et al.
Pubblicazione: (2025)
Scaling Open-Ended Reasoning to Predict the Future
di: Chandak, Nikhil, et al.
Pubblicazione: (2025)
di: Chandak, Nikhil, et al.
Pubblicazione: (2025)
Visual Preference Optimization with Rubric Rewards
di: Yu, Ya-Qi, et al.
Pubblicazione: (2026)
di: Yu, Ya-Qi, et al.
Pubblicazione: (2026)
Reinforcement Learning with Robust Rubric Rewards
di: Yu, Ya-Qi, et al.
Pubblicazione: (2026)
di: Yu, Ya-Qi, et al.
Pubblicazione: (2026)
No Regrets: Investigating and Improving Regret Approximations for Curriculum Discovery
di: Rutherford, Alexander, et al.
Pubblicazione: (2024)
di: Rutherford, Alexander, et al.
Pubblicazione: (2024)
Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling
di: Xie, Lipeng, et al.
Pubblicazione: (2025)
di: Xie, Lipeng, et al.
Pubblicazione: (2025)
Reward and Guidance through Rubrics: Promoting Exploration to Improve Multi-Domain Reasoning
di: Bi, Baolong, et al.
Pubblicazione: (2025)
di: Bi, Baolong, et al.
Pubblicazione: (2025)
Doubly-Robust LLM-as-a-Judge: Externally Valid Estimation with Imperfect Personas
di: Guerdan, Luke, et al.
Pubblicazione: (2025)
di: Guerdan, Luke, et al.
Pubblicazione: (2025)
Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA
di: Belmadani, Ikram, et al.
Pubblicazione: (2026)
di: Belmadani, Ikram, et al.
Pubblicazione: (2026)
Learning to Judge: LLMs Designing and Applying Evaluation Rubrics
di: Siro, Clemencia, et al.
Pubblicazione: (2026)
di: Siro, Clemencia, et al.
Pubblicazione: (2026)
How can representation dimension dominate structurally pruned LLMs?
di: Xu, Mingxue, et al.
Pubblicazione: (2025)
di: Xu, Mingxue, et al.
Pubblicazione: (2025)
RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards
di: Li, Gaotang, et al.
Pubblicazione: (2026)
di: Li, Gaotang, et al.
Pubblicazione: (2026)
Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
di: Wu, Tianhao, et al.
Pubblicazione: (2024)
di: Wu, Tianhao, et al.
Pubblicazione: (2024)
PAnDA: Rethinking Metric Differential Privacy Optimization at Scale with Anchor-Based Approximation
di: Liu, Ruiyao, et al.
Pubblicazione: (2025)
di: Liu, Ruiyao, et al.
Pubblicazione: (2025)
El diccionario
di: Laura Chenillo Alazraki
Pubblicazione: (2013)
di: Laura Chenillo Alazraki
Pubblicazione: (2013)
Documenti analoghi
-
Training AI Co-Scientists Using Rubric Rewards
di: Goel, Shashwat, et al.
Pubblicazione: (2025) -
Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
di: Collot, Stephane, et al.
Pubblicazione: (2025) -
Evaluating Privacy Leakage in Split Learning
di: Qiu, Xinchi, et al.
Pubblicazione: (2023) -
Scaling Small Agents Through Strategy Auctions
di: Alazraki, Lisa, et al.
Pubblicazione: (2026) -
Meta-Reasoning Improves Tool Use in Large Language Models
di: Alazraki, Lisa, et al.
Pubblicazione: (2024)