Rethinking Rubric Generation for Improving LLM Judge and Reward Modeling for Open-ended Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shen, William F., Qiu, Xinchi, Whitehouse, Chenxi, Alazraki, Lisa, Goel, Shashwat, Barbieri, Francesco, Willi, Timon, Mathur, Akhil, Leontiadis, Ilias |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Training AI Co-Scientists Using Rubric Rewards
von: Goel, Shashwat, et al.
Veröffentlicht: (2025)
von: Goel, Shashwat, et al.
Veröffentlicht: (2025)
Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
von: Collot, Stephane, et al.
Veröffentlicht: (2025)
von: Collot, Stephane, et al.
Veröffentlicht: (2025)
Evaluating Privacy Leakage in Split Learning
von: Qiu, Xinchi, et al.
Veröffentlicht: (2023)
von: Qiu, Xinchi, et al.
Veröffentlicht: (2023)
Scaling Small Agents Through Strategy Auctions
von: Alazraki, Lisa, et al.
Veröffentlicht: (2026)
von: Alazraki, Lisa, et al.
Veröffentlicht: (2026)
Meta-Reasoning Improves Tool Use in Large Language Models
von: Alazraki, Lisa, et al.
Veröffentlicht: (2024)
von: Alazraki, Lisa, et al.
Veröffentlicht: (2024)
AdaRubric: Task-Adaptive Rubrics for Reliable LLM Agent Evaluation and Reward Learning
von: Ding, Liang
Veröffentlicht: (2026)
von: Ding, Liang
Veröffentlicht: (2026)
OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM Alignment
von: Liu, Tianci, et al.
Veröffentlicht: (2025)
von: Liu, Tianci, et al.
Veröffentlicht: (2025)
Culturally Grounded Physical Commonsense Reasoning in Italian and English: A Submission to the MRL 2025 Shared Task
von: De Santis, Marco, et al.
Veröffentlicht: (2025)
von: De Santis, Marco, et al.
Veröffentlicht: (2025)
Towards Knowledge-Grounded Natural Language Understanding and Generation
von: Whitehouse, Chenxi
Veröffentlicht: (2024)
von: Whitehouse, Chenxi
Veröffentlicht: (2024)
Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation
von: Radharapu, Bhaktipriya, et al.
Veröffentlicht: (2025)
von: Radharapu, Bhaktipriya, et al.
Veröffentlicht: (2025)
Rubric-Grounded RL: Structured Judge Rewards for Generalizable Reasoning
von: Bhattarai, Manish, et al.
Veröffentlicht: (2026)
von: Bhattarai, Manish, et al.
Veröffentlicht: (2026)
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
von: Whitehouse, Chenxi, et al.
Veröffentlicht: (2025)
von: Whitehouse, Chenxi, et al.
Veröffentlicht: (2025)
Step-wise Rubric Rewards for LLM Reasoning
von: Xie, Weichu, et al.
Veröffentlicht: (2026)
von: Xie, Weichu, et al.
Veröffentlicht: (2026)
RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following
von: Pan, Tianjun, et al.
Veröffentlicht: (2026)
von: Pan, Tianjun, et al.
Veröffentlicht: (2026)
Rubrics to Tokens: Bridging Response-level Rubrics and Token-level Rewards in Instruction Following Tasks
von: Xu, Tianze, et al.
Veröffentlicht: (2026)
von: Xu, Tianze, et al.
Veröffentlicht: (2026)
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
von: Ding, Ruomeng, et al.
Veröffentlicht: (2026)
von: Ding, Ruomeng, et al.
Veröffentlicht: (2026)
Enhancing LLM Robustness to Perturbed Instructions: An Empirical Study
von: Agrawal, Aryan, et al.
Veröffentlicht: (2025)
von: Agrawal, Aryan, et al.
Veröffentlicht: (2025)
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind
von: Lupu, Andrei, et al.
Veröffentlicht: (2025)
von: Lupu, Andrei, et al.
Veröffentlicht: (2025)
Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards
von: Yuan, Youliang, et al.
Veröffentlicht: (2025)
von: Yuan, Youliang, et al.
Veröffentlicht: (2025)
Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text
von: Huang, Chengyu, et al.
Veröffentlicht: (2026)
von: Huang, Chengyu, et al.
Veröffentlicht: (2026)
Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning
von: Ye, Zhiling, et al.
Veröffentlicht: (2025)
von: Ye, Zhiling, et al.
Veröffentlicht: (2025)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
von: Hong, Yihan, et al.
Veröffentlicht: (2026)
von: Hong, Yihan, et al.
Veröffentlicht: (2026)
Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR
von: Tyagi, Utkarsh, et al.
Veröffentlicht: (2026)
von: Tyagi, Utkarsh, et al.
Veröffentlicht: (2026)
AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning
von: Jia, Mengzhao, et al.
Veröffentlicht: (2025)
von: Jia, Mengzhao, et al.
Veröffentlicht: (2025)
A Judge-free LLM Open-ended Generation Benchmark Based on the Distributional Hypothesis
von: Imajo, Kentaro, et al.
Veröffentlicht: (2025)
von: Imajo, Kentaro, et al.
Veröffentlicht: (2025)
ProBench: Judging Multimodal Foundation Models on Open-ended Multi-domain Expert Tasks
von: Yang, Yan, et al.
Veröffentlicht: (2025)
von: Yang, Yan, et al.
Veröffentlicht: (2025)
Scaling Open-Ended Reasoning to Predict the Future
von: Chandak, Nikhil, et al.
Veröffentlicht: (2025)
von: Chandak, Nikhil, et al.
Veröffentlicht: (2025)
Visual Preference Optimization with Rubric Rewards
von: Yu, Ya-Qi, et al.
Veröffentlicht: (2026)
von: Yu, Ya-Qi, et al.
Veröffentlicht: (2026)
Reinforcement Learning with Robust Rubric Rewards
von: Yu, Ya-Qi, et al.
Veröffentlicht: (2026)
von: Yu, Ya-Qi, et al.
Veröffentlicht: (2026)
No Regrets: Investigating and Improving Regret Approximations for Curriculum Discovery
von: Rutherford, Alexander, et al.
Veröffentlicht: (2024)
von: Rutherford, Alexander, et al.
Veröffentlicht: (2024)
Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling
von: Xie, Lipeng, et al.
Veröffentlicht: (2025)
von: Xie, Lipeng, et al.
Veröffentlicht: (2025)
Reward and Guidance through Rubrics: Promoting Exploration to Improve Multi-Domain Reasoning
von: Bi, Baolong, et al.
Veröffentlicht: (2025)
von: Bi, Baolong, et al.
Veröffentlicht: (2025)
Doubly-Robust LLM-as-a-Judge: Externally Valid Estimation with Imperfect Personas
von: Guerdan, Luke, et al.
Veröffentlicht: (2025)
von: Guerdan, Luke, et al.
Veröffentlicht: (2025)
Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA
von: Belmadani, Ikram, et al.
Veröffentlicht: (2026)
von: Belmadani, Ikram, et al.
Veröffentlicht: (2026)
Learning to Judge: LLMs Designing and Applying Evaluation Rubrics
von: Siro, Clemencia, et al.
Veröffentlicht: (2026)
von: Siro, Clemencia, et al.
Veröffentlicht: (2026)
How can representation dimension dominate structurally pruned LLMs?
von: Xu, Mingxue, et al.
Veröffentlicht: (2025)
von: Xu, Mingxue, et al.
Veröffentlicht: (2025)
RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards
von: Li, Gaotang, et al.
Veröffentlicht: (2026)
von: Li, Gaotang, et al.
Veröffentlicht: (2026)
Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
PAnDA: Rethinking Metric Differential Privacy Optimization at Scale with Anchor-Based Approximation
von: Liu, Ruiyao, et al.
Veröffentlicht: (2025)
von: Liu, Ruiyao, et al.
Veröffentlicht: (2025)
El diccionario
von: Laura Chenillo Alazraki
Veröffentlicht: (2013)
von: Laura Chenillo Alazraki
Veröffentlicht: (2013)
Ähnliche Einträge
-
Training AI Co-Scientists Using Rubric Rewards
von: Goel, Shashwat, et al.
Veröffentlicht: (2025) -
Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
von: Collot, Stephane, et al.
Veröffentlicht: (2025) -
Evaluating Privacy Leakage in Split Learning
von: Qiu, Xinchi, et al.
Veröffentlicht: (2023) -
Scaling Small Agents Through Strategy Auctions
von: Alazraki, Lisa, et al.
Veröffentlicht: (2026) -
Meta-Reasoning Improves Tool Use in Large Language Models
von: Alazraki, Lisa, et al.
Veröffentlicht: (2024)