Guardado en:
| Autores principales: | Shen, William F., Qiu, Xinchi, Whitehouse, Chenxi, Alazraki, Lisa, Goel, Shashwat, Barbieri, Francesco, Willi, Timon, Mathur, Akhil, Leontiadis, Ilias |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2602.05125 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Training AI Co-Scientists Using Rubric Rewards
por: Goel, Shashwat, et al.
Publicado: (2025)
por: Goel, Shashwat, et al.
Publicado: (2025)
Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
por: Collot, Stephane, et al.
Publicado: (2025)
por: Collot, Stephane, et al.
Publicado: (2025)
Evaluating Privacy Leakage in Split Learning
por: Qiu, Xinchi, et al.
Publicado: (2023)
por: Qiu, Xinchi, et al.
Publicado: (2023)
Scaling Small Agents Through Strategy Auctions
por: Alazraki, Lisa, et al.
Publicado: (2026)
por: Alazraki, Lisa, et al.
Publicado: (2026)
Meta-Reasoning Improves Tool Use in Large Language Models
por: Alazraki, Lisa, et al.
Publicado: (2024)
por: Alazraki, Lisa, et al.
Publicado: (2024)
Culturally Grounded Physical Commonsense Reasoning in Italian and English: A Submission to the MRL 2025 Shared Task
por: De Santis, Marco, et al.
Publicado: (2025)
por: De Santis, Marco, et al.
Publicado: (2025)
AdaRubric: Task-Adaptive Rubrics for Reliable LLM Agent Evaluation and Reward Learning
por: Ding, Liang
Publicado: (2026)
por: Ding, Liang
Publicado: (2026)
OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM Alignment
por: Liu, Tianci, et al.
Publicado: (2025)
por: Liu, Tianci, et al.
Publicado: (2025)
Towards Knowledge-Grounded Natural Language Understanding and Generation
por: Whitehouse, Chenxi
Publicado: (2024)
por: Whitehouse, Chenxi
Publicado: (2024)
Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation
por: Radharapu, Bhaktipriya, et al.
Publicado: (2025)
por: Radharapu, Bhaktipriya, et al.
Publicado: (2025)
Rubric-Grounded RL: Structured Judge Rewards for Generalizable Reasoning
por: Bhattarai, Manish, et al.
Publicado: (2026)
por: Bhattarai, Manish, et al.
Publicado: (2026)
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
por: Whitehouse, Chenxi, et al.
Publicado: (2025)
por: Whitehouse, Chenxi, et al.
Publicado: (2025)
Enhancing LLM Robustness to Perturbed Instructions: An Empirical Study
por: Agrawal, Aryan, et al.
Publicado: (2025)
por: Agrawal, Aryan, et al.
Publicado: (2025)
Step-wise Rubric Rewards for LLM Reasoning
por: Xie, Weichu, et al.
Publicado: (2026)
por: Xie, Weichu, et al.
Publicado: (2026)
RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following
por: Pan, Tianjun, et al.
Publicado: (2026)
por: Pan, Tianjun, et al.
Publicado: (2026)
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind
por: Lupu, Andrei, et al.
Publicado: (2025)
por: Lupu, Andrei, et al.
Publicado: (2025)
Rubrics to Tokens: Bridging Response-level Rubrics and Token-level Rewards in Instruction Following Tasks
por: Xu, Tianze, et al.
Publicado: (2026)
por: Xu, Tianze, et al.
Publicado: (2026)
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
por: Ding, Ruomeng, et al.
Publicado: (2026)
por: Ding, Ruomeng, et al.
Publicado: (2026)
Scaling Open-Ended Reasoning to Predict the Future
por: Chandak, Nikhil, et al.
Publicado: (2025)
por: Chandak, Nikhil, et al.
Publicado: (2025)
Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards
por: Yuan, Youliang, et al.
Publicado: (2025)
por: Yuan, Youliang, et al.
Publicado: (2025)
No Regrets: Investigating and Improving Regret Approximations for Curriculum Discovery
por: Rutherford, Alexander, et al.
Publicado: (2024)
por: Rutherford, Alexander, et al.
Publicado: (2024)
Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning
por: Ye, Zhiling, et al.
Publicado: (2025)
por: Ye, Zhiling, et al.
Publicado: (2025)
Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text
por: Huang, Chengyu, et al.
Publicado: (2026)
por: Huang, Chengyu, et al.
Publicado: (2026)
How can representation dimension dominate structurally pruned LLMs?
por: Xu, Mingxue, et al.
Publicado: (2025)
por: Xu, Mingxue, et al.
Publicado: (2025)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
por: Hong, Yihan, et al.
Publicado: (2026)
por: Hong, Yihan, et al.
Publicado: (2026)
Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR
por: Tyagi, Utkarsh, et al.
Publicado: (2026)
por: Tyagi, Utkarsh, et al.
Publicado: (2026)
AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning
por: Jia, Mengzhao, et al.
Publicado: (2025)
por: Jia, Mengzhao, et al.
Publicado: (2025)
El Archivo Histórico del Instituto Nacional de Migración
por: Paola Chenillo Alazraki
Publicado: (2008)
por: Paola Chenillo Alazraki
Publicado: (2008)
El diccionario
por: Laura Chenillo Alazraki
Publicado: (2013)
por: Laura Chenillo Alazraki
Publicado: (2013)
Doubly-Robust LLM-as-a-Judge: Externally Valid Estimation with Imperfect Personas
por: Guerdan, Luke, et al.
Publicado: (2025)
por: Guerdan, Luke, et al.
Publicado: (2025)
Proportional Aggregation of Preferences for Sequential Decision Making
por: Chandak, Nikhil, et al.
Publicado: (2023)
por: Chandak, Nikhil, et al.
Publicado: (2023)
ProBench: Judging Multimodal Foundation Models on Open-ended Multi-domain Expert Tasks
por: Yang, Yan, et al.
Publicado: (2025)
por: Yang, Yan, et al.
Publicado: (2025)
A Judge-free LLM Open-ended Generation Benchmark Based on the Distributional Hypothesis
por: Imajo, Kentaro, et al.
Publicado: (2025)
por: Imajo, Kentaro, et al.
Publicado: (2025)
Visual Preference Optimization with Rubric Rewards
por: Yu, Ya-Qi, et al.
Publicado: (2026)
por: Yu, Ya-Qi, et al.
Publicado: (2026)
Reinforcement Learning with Robust Rubric Rewards
por: Yu, Ya-Qi, et al.
Publicado: (2026)
por: Yu, Ya-Qi, et al.
Publicado: (2026)
Improving the OOD Performance of Closed-Source LLMs on NLI Through Strategic Data Selection
por: Stacey, Joe, et al.
Publicado: (2025)
por: Stacey, Joe, et al.
Publicado: (2025)
Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling
por: Xie, Lipeng, et al.
Publicado: (2025)
por: Xie, Lipeng, et al.
Publicado: (2025)
Reward and Guidance through Rubrics: Promoting Exploration to Improve Multi-Domain Reasoning
por: Bi, Baolong, et al.
Publicado: (2025)
por: Bi, Baolong, et al.
Publicado: (2025)
PAnDA: Rethinking Metric Differential Privacy Optimization at Scale with Anchor-Based Approximation
por: Liu, Ruiyao, et al.
Publicado: (2025)
por: Liu, Ruiyao, et al.
Publicado: (2025)
Learning to Judge: LLMs Designing and Applying Evaluation Rubrics
por: Siro, Clemencia, et al.
Publicado: (2026)
por: Siro, Clemencia, et al.
Publicado: (2026)
Ejemplares similares
-
Training AI Co-Scientists Using Rubric Rewards
por: Goel, Shashwat, et al.
Publicado: (2025) -
Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
por: Collot, Stephane, et al.
Publicado: (2025) -
Evaluating Privacy Leakage in Split Learning
por: Qiu, Xinchi, et al.
Publicado: (2023) -
Scaling Small Agents Through Strategy Auctions
por: Alazraki, Lisa, et al.
Publicado: (2026) -
Meta-Reasoning Improves Tool Use in Large Language Models
por: Alazraki, Lisa, et al.
Publicado: (2024)