Quantifying the Statistical Effect of Rubric Modifications on Human-Autorater Agreement
Fuente:
arXiv
Guardado en:
| Autores principales: | Huynh, Jessica, Gomez, Alfredo, Deviyani, Athiya, Shelby, Renee, Bigham, Jeffrey P., Diaz, Fernando |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Contextual Metric Meta-Evaluation by Measuring Local Metric Accuracy
por: Deviyani, Athiya, et al.
Publicado: (2025)
por: Deviyani, Athiya, et al.
Publicado: (2025)
Taxonomy of User Needs and Actions
por: Shelby, Renee, et al.
Publicado: (2025)
por: Shelby, Renee, et al.
Publicado: (2025)
Judging with Confidence: Calibrating Autoraters to Preference Distributions
por: Li, Zhuohang, et al.
Publicado: (2025)
por: Li, Zhuohang, et al.
Publicado: (2025)
Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation
por: Vu, Tu, et al.
Publicado: (2024)
por: Vu, Tu, et al.
Publicado: (2024)
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set
por: Finkelstein, Mara, et al.
Publicado: (2024)
por: Finkelstein, Mara, et al.
Publicado: (2024)
DesignPref: Capturing Personal Preferences in Visual Design Generation
por: Peng, Yi-Hao, et al.
Publicado: (2025)
por: Peng, Yi-Hao, et al.
Publicado: (2025)
Morae: Proactively Pausing UI Agents for User Choices
por: Peng, Yi-Hao, et al.
Publicado: (2025)
por: Peng, Yi-Hao, et al.
Publicado: (2025)
Nonparametric LLM Evaluation from Preference Data
por: Frauen, Dennis, et al.
Publicado: (2026)
por: Frauen, Dennis, et al.
Publicado: (2026)
UICoder: Finetuning Large Language Models to Generate User Interface Code through Automated Feedback
por: Wu, Jason, et al.
Publicado: (2024)
por: Wu, Jason, et al.
Publicado: (2024)
Open Rubric System: Scaling Reinforcement Learning with Pairwise Adaptive Rubric
por: Jia, Ruipeng, et al.
Publicado: (2026)
por: Jia, Ruipeng, et al.
Publicado: (2026)
AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning
por: Jia, Mengzhao, et al.
Publicado: (2025)
por: Jia, Mengzhao, et al.
Publicado: (2025)
Modeling Distinct Human Interaction in Web Agents
por: Huq, Faria, et al.
Publicado: (2026)
por: Huq, Faria, et al.
Publicado: (2026)
CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation
por: Huq, Faria, et al.
Publicado: (2025)
por: Huq, Faria, et al.
Publicado: (2025)
EvoRubric: Self-Evolving Rubric-Driven RL for Open-Ended Generation
por: Guan, Xin, et al.
Publicado: (2026)
por: Guan, Xin, et al.
Publicado: (2026)
UIClip: A Data-driven Model for Assessing User Interface Design
por: Wu, Jason, et al.
Publicado: (2024)
por: Wu, Jason, et al.
Publicado: (2024)
OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM Alignment
por: Liu, Tianci, et al.
Publicado: (2025)
por: Liu, Tianci, et al.
Publicado: (2025)
HIDAgent: A Toolkit Enabling "Personal Agents" on HID-Compatible Devices
por: Bigham, Jeffrey P.
Publicado: (2026)
por: Bigham, Jeffrey P.
Publicado: (2026)
Diverging Transformer Predictions for Human Sentence Processing: A Comprehensive Analysis of Agreement Attraction Effects
por: von der Malsburg, Titus, et al.
Publicado: (2026)
por: von der Malsburg, Titus, et al.
Publicado: (2026)
Learning Query-Specific Rubrics from Human Preferences for DeepResearch Report Generation
por: Lv, Changze, et al.
Publicado: (2026)
por: Lv, Changze, et al.
Publicado: (2026)
Deep Research as Rubric for Reinforcement Learning
por: Mei, Wangyi, et al.
Publicado: (2026)
por: Mei, Wangyi, et al.
Publicado: (2026)
RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards
por: Li, Gaotang, et al.
Publicado: (2026)
por: Li, Gaotang, et al.
Publicado: (2026)
AdaRubric: Task-Adaptive Rubrics for Reliable LLM Agent Evaluation and Reward Learning
por: Ding, Liang
Publicado: (2026)
por: Ding, Liang
Publicado: (2026)
Case-Specific Rubrics for Clinical AI Evaluation: Methodology, Validation, and LLM-Clinician Agreement Across 823 Encounters
por: Shah, Aaryan, et al.
Publicado: (2026)
por: Shah, Aaryan, et al.
Publicado: (2026)
Preference-Aware Rubric Learning for Personalized Evaluation
por: Qiu, Yilun, et al.
Publicado: (2026)
por: Qiu, Yilun, et al.
Publicado: (2026)
A Comprehensive Rubric for Annotating Pathological Speech
por: Corrales-Astorgano, Mario, et al.
Publicado: (2024)
por: Corrales-Astorgano, Mario, et al.
Publicado: (2024)
Rubrics to Tokens: Bridging Response-level Rubrics and Token-level Rewards in Instruction Following Tasks
por: Xu, Tianze, et al.
Publicado: (2026)
por: Xu, Tianze, et al.
Publicado: (2026)
GenAudit: Fixing Factual Errors in Language Model Outputs with Evidence
por: Krishna, Kundan, et al.
Publicado: (2024)
por: Krishna, Kundan, et al.
Publicado: (2024)
Debiasing Text Safety Classifiers through a Fairness-Aware Ensemble
por: Sturman, Olivia, et al.
Publicado: (2024)
por: Sturman, Olivia, et al.
Publicado: (2024)
LLM Essay Scoring Under Holistic and Analytic Rubrics: Prompt Effects and Bias
por: Kucia, Filip J., et al.
Publicado: (2026)
por: Kucia, Filip J., et al.
Publicado: (2026)
A Unified Framework to Quantify Cultural Intelligence of AI
por: Dev, Sunipa, et al.
Publicado: (2026)
por: Dev, Sunipa, et al.
Publicado: (2026)
ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents
por: Sharma, Manasi, et al.
Publicado: (2025)
por: Sharma, Manasi, et al.
Publicado: (2025)
Quantifying and Predicting Disagreement in Graded Human Ratings
por: Zhang, Leixin, et al.
Publicado: (2026)
por: Zhang, Leixin, et al.
Publicado: (2026)
Policy Maps: Tools for Guiding the Unbounded Space of LLM Behaviors
por: Lam, Michelle S., et al.
Publicado: (2024)
por: Lam, Michelle S., et al.
Publicado: (2024)
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
por: Jung, Jaehun, et al.
Publicado: (2024)
por: Jung, Jaehun, et al.
Publicado: (2024)
AMARIS: A Memory-Augmented Rubric Improvement System for Rubric-Based Reinforcement Learning
por: Wu, Peilin, et al.
Publicado: (2026)
por: Wu, Peilin, et al.
Publicado: (2026)
ARES: Automated Rubric Synthesis for Scalable LLM Reinforcement Learning
por: Li, Xiaoyuan, et al.
Publicado: (2026)
por: Li, Xiaoyuan, et al.
Publicado: (2026)
Think-with-Rubrics: From External Evaluator to Internal Reasoning Guidance
por: Yu, Jiachen, et al.
Publicado: (2026)
por: Yu, Jiachen, et al.
Publicado: (2026)
Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards
por: Yuan, Youliang, et al.
Publicado: (2025)
por: Yuan, Youliang, et al.
Publicado: (2025)
Modelling Adjectival Modification Effects on Semantic Plausibility
por: Golub, Anna, et al.
Publicado: (2025)
por: Golub, Anna, et al.
Publicado: (2025)
Reinforcement Learning with Rubric Anchors
por: Huang, Zenan, et al.
Publicado: (2025)
por: Huang, Zenan, et al.
Publicado: (2025)
Ejemplares similares
-
Contextual Metric Meta-Evaluation by Measuring Local Metric Accuracy
por: Deviyani, Athiya, et al.
Publicado: (2025) -
Taxonomy of User Needs and Actions
por: Shelby, Renee, et al.
Publicado: (2025) -
Judging with Confidence: Calibrating Autoraters to Preference Distributions
por: Li, Zhuohang, et al.
Publicado: (2025) -
Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation
por: Vu, Tu, et al.
Publicado: (2024) -
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set
por: Finkelstein, Mara, et al.
Publicado: (2024)