Rubric-Grounded RL: Structured Judge Rewards for Generalizable Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | Bhattarai, Manish, Boureima, Ismael, Ranasinghe, Nishath Rajiv, Pakin, Scott, O'Malley, Dan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Benchmarking Large Language Models with Integer Sequence Generation Tasks
di: O'Malley, Daniel, et al.
Pubblicazione: (2024)
di: O'Malley, Daniel, et al.
Pubblicazione: (2024)
Lost in OCR Translation? Vision-Based Approaches to Robust Document Retrieval
di: Most, Alexander, et al.
Pubblicazione: (2025)
di: Most, Alexander, et al.
Pubblicazione: (2025)
ARCS: Agentic Retrieval-Augmented Code Synthesis with Iterative Refinement
di: Bhattarai, Manish, et al.
Pubblicazione: (2025)
di: Bhattarai, Manish, et al.
Pubblicazione: (2025)
Enhancing Cross-Language Code Translation via Task-Specific Embedding Alignment in Retrieval-Augmented Generation
di: Bhattarai, Manish, et al.
Pubblicazione: (2024)
di: Bhattarai, Manish, et al.
Pubblicazione: (2024)
Tensor Train Low-rank Approximation (TT-LoRA): Democratizing AI with Accelerated LLMs
di: Anjum, Afia, et al.
Pubblicazione: (2024)
di: Anjum, Afia, et al.
Pubblicazione: (2024)
Enhancing Code Translation in Language Models with Few-Shot Learning via Retrieval-Augmented Generation
di: Bhattarai, Manish, et al.
Pubblicazione: (2024)
di: Bhattarai, Manish, et al.
Pubblicazione: (2024)
LLM-Assisted Translation of Legacy FORTRAN Codes to C++: A Cross-Platform Study
di: Ranasinghe, Nishath Rajiv, et al.
Pubblicazione: (2025)
di: Ranasinghe, Nishath Rajiv, et al.
Pubblicazione: (2025)
Binary Bleed: Fast Distributed and Parallel Method for Automatic Model Selection
di: Barron, Ryan, et al.
Pubblicazione: (2024)
di: Barron, Ryan, et al.
Pubblicazione: (2024)
Rethinking Rubric Generation for Improving LLM Judge and Reward Modeling for Open-ended Tasks
di: Shen, William F., et al.
Pubblicazione: (2026)
di: Shen, William F., et al.
Pubblicazione: (2026)
Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR
di: Tyagi, Utkarsh, et al.
Pubblicazione: (2026)
di: Tyagi, Utkarsh, et al.
Pubblicazione: (2026)
Reward and Guidance through Rubrics: Promoting Exploration to Improve Multi-Domain Reasoning
di: Bi, Baolong, et al.
Pubblicazione: (2025)
di: Bi, Baolong, et al.
Pubblicazione: (2025)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
di: Hong, Yihan, et al.
Pubblicazione: (2026)
di: Hong, Yihan, et al.
Pubblicazione: (2026)
RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following
di: Pan, Tianjun, et al.
Pubblicazione: (2026)
di: Pan, Tianjun, et al.
Pubblicazione: (2026)
Generating Data-Driven Reasoning Rubrics for Domain-Adaptive Reward Modeling
di: Sanders, Kate, et al.
Pubblicazione: (2026)
di: Sanders, Kate, et al.
Pubblicazione: (2026)
HEAL: Hierarchical Embedding Alignment Loss for Improved Retrieval and Representation Learning
di: Bhattarai, Manish, et al.
Pubblicazione: (2024)
di: Bhattarai, Manish, et al.
Pubblicazione: (2024)
Reward Hacking in Rubric-Based Reinforcement Learning
di: Mahmoud, Anas, et al.
Pubblicazione: (2026)
di: Mahmoud, Anas, et al.
Pubblicazione: (2026)
Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling
di: Xie, Lipeng, et al.
Pubblicazione: (2025)
di: Xie, Lipeng, et al.
Pubblicazione: (2025)
mR3: Multilingual Rubric-Agnostic Reward Reasoning Models
di: Anugraha, David, et al.
Pubblicazione: (2025)
di: Anugraha, David, et al.
Pubblicazione: (2025)
Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning
di: Ye, Zhiling, et al.
Pubblicazione: (2025)
di: Ye, Zhiling, et al.
Pubblicazione: (2025)
Do LLMs Judge Distantly Supervised Named Entity Labels Well? Constructing the JudgeWEL Dataset
di: Plum, Alistair, et al.
Pubblicazione: (2026)
di: Plum, Alistair, et al.
Pubblicazione: (2026)
AdaRubric: Task-Adaptive Rubrics for Reliable LLM Agent Evaluation and Reward Learning
di: Ding, Liang
Pubblicazione: (2026)
di: Ding, Liang
Pubblicazione: (2026)
Visual Preference Optimization with Rubric Rewards
di: Yu, Ya-Qi, et al.
Pubblicazione: (2026)
di: Yu, Ya-Qi, et al.
Pubblicazione: (2026)
Reinforcement Learning with Robust Rubric Rewards
di: Yu, Ya-Qi, et al.
Pubblicazione: (2026)
di: Yu, Ya-Qi, et al.
Pubblicazione: (2026)
Rubrics to Tokens: Bridging Response-level Rubrics and Token-level Rewards in Instruction Following Tasks
di: Xu, Tianze, et al.
Pubblicazione: (2026)
di: Xu, Tianze, et al.
Pubblicazione: (2026)
When Rubrics Fail: Error Enumeration as Reward in Reference-Free RL Post-Training for Virtual Try-On
di: Ikezogwo, Wisdom, et al.
Pubblicazione: (2026)
di: Ikezogwo, Wisdom, et al.
Pubblicazione: (2026)
Rubric-Guided Process Reward for Stepwise Model Routing
di: Ye, Shenghao, et al.
Pubblicazione: (2026)
di: Ye, Shenghao, et al.
Pubblicazione: (2026)
The Transparent Earth: A Multimodal Foundation Model for the Earth's Subsurface
di: Mazumder, Arnab, et al.
Pubblicazione: (2025)
di: Mazumder, Arnab, et al.
Pubblicazione: (2025)
Approximately Optimal Search on a Higher-dimensional Sliding Puzzle
di: Merleau, Nono SC, et al.
Pubblicazione: (2024)
di: Merleau, Nono SC, et al.
Pubblicazione: (2024)
Retrieval-Augmented Process Reward Model for Generalizable Mathematical Reasoning
di: Zhu, Jiachen, et al.
Pubblicazione: (2025)
di: Zhu, Jiachen, et al.
Pubblicazione: (2025)
Auto-Rubric as Reward: From Implicit Preferences to Explicit Multimodal Generative Criteria
di: Tian, Juanxi, et al.
Pubblicazione: (2026)
di: Tian, Juanxi, et al.
Pubblicazione: (2026)
MTRE: Multi-Token Reliability Estimation for Hallucination Detection in VLMs
di: Zollicoffer, Geigh, et al.
Pubblicazione: (2025)
di: Zollicoffer, Geigh, et al.
Pubblicazione: (2025)
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
di: Ding, Ruomeng, et al.
Pubblicazione: (2026)
di: Ding, Ruomeng, et al.
Pubblicazione: (2026)
FlowRL: Matching Reward Distributions for LLM Reasoning
di: Zhu, Xuekai, et al.
Pubblicazione: (2025)
di: Zhu, Xuekai, et al.
Pubblicazione: (2025)
R3: Robust Rubric-Agnostic Reward Models
di: Anugraha, David, et al.
Pubblicazione: (2025)
di: Anugraha, David, et al.
Pubblicazione: (2025)
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
di: Gunjal, Anisha, et al.
Pubblicazione: (2025)
di: Gunjal, Anisha, et al.
Pubblicazione: (2025)
Kardia-R1: Unleashing LLMs to Reason toward Understanding and Empathy for Emotional Support via Rubric-as-Judge Reinforcement Learning
di: Yuan, Jiahao, et al.
Pubblicazione: (2025)
di: Yuan, Jiahao, et al.
Pubblicazione: (2025)
Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning
di: Mahrooghi, Ilia, et al.
Pubblicazione: (2026)
di: Mahrooghi, Ilia, et al.
Pubblicazione: (2026)
Multidimensional Rubric-oriented Reward Model Learning via Geometric Projection Reference Constraints
di: Jin, Yongnan, et al.
Pubblicazione: (2025)
di: Jin, Yongnan, et al.
Pubblicazione: (2025)
CDRRM: Contrast-Driven Rubric Generation for Reliable and Interpretable Reward Modeling
di: Liu, Dengcan, et al.
Pubblicazione: (2026)
di: Liu, Dengcan, et al.
Pubblicazione: (2026)
CrystalReasoner: Reasoning and RL for Property-Conditioned Crystal Structure Generation
di: Wu, Yuyang, et al.
Pubblicazione: (2026)
di: Wu, Yuyang, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Benchmarking Large Language Models with Integer Sequence Generation Tasks
di: O'Malley, Daniel, et al.
Pubblicazione: (2024) -
Lost in OCR Translation? Vision-Based Approaches to Robust Document Retrieval
di: Most, Alexander, et al.
Pubblicazione: (2025) -
ARCS: Agentic Retrieval-Augmented Code Synthesis with Iterative Refinement
di: Bhattarai, Manish, et al.
Pubblicazione: (2025) -
Enhancing Cross-Language Code Translation via Task-Specific Embedding Alignment in Retrieval-Augmented Generation
di: Bhattarai, Manish, et al.
Pubblicazione: (2024) -
Tensor Train Low-rank Approximation (TT-LoRA): Democratizing AI with Accelerated LLMs
di: Anjum, Afia, et al.
Pubblicazione: (2024)