Approximating Human Preferences Using a Multi-Judge Learned System
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908619789828096 |
|---|---|
| author | Sprejer, Eitán Avalos, Fernando Bernardi, Augusto Faustino, Jose Pedro Brito de Azevedo Haimes, Jacob Oozeer, Narmeen Fatimah |
| author_facet | Sprejer, Eitán Avalos, Fernando Bernardi, Augusto Faustino, Jose Pedro Brito de Azevedo Haimes, Jacob Oozeer, Narmeen Fatimah |
| contents | Aligning LLM-based judges with human preferences is a significant challenge, as they are difficult to calibrate and often suffer from rubric sensitivity, bias, and instability. Overcoming this challenge advances key applications, such as creating reliable reward models for Reinforcement Learning from Human Feedback (RLHF) and building effective routing systems that select the best-suited model for a given user query. In this work, we propose a framework for modeling diverse, persona-based preferences by learning to aggregate outputs from multiple rubric-conditioned judges. We investigate the performance of this approach against naive baselines and assess its robustness through case studies on both human and LLM-judges biases. Our primary contributions include a persona-based method for synthesizing preference labels at scale and two distinct implementations of our aggregator: Generalized Additive Model (GAM) and a Multi-Layer Perceptron (MLP). |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_25884 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Approximating Human Preferences Using a Multi-Judge Learned System Sprejer, Eitán Avalos, Fernando Bernardi, Augusto Faustino, Jose Pedro Brito de Azevedo Haimes, Jacob Oozeer, Narmeen Fatimah Artificial Intelligence Computation and Language Machine Learning Aligning LLM-based judges with human preferences is a significant challenge, as they are difficult to calibrate and often suffer from rubric sensitivity, bias, and instability. Overcoming this challenge advances key applications, such as creating reliable reward models for Reinforcement Learning from Human Feedback (RLHF) and building effective routing systems that select the best-suited model for a given user query. In this work, we propose a framework for modeling diverse, persona-based preferences by learning to aggregate outputs from multiple rubric-conditioned judges. We investigate the performance of this approach against naive baselines and assess its robustness through case studies on both human and LLM-judges biases. Our primary contributions include a persona-based method for synthesizing preference labels at scale and two distinct implementations of our aggregator: Generalized Additive Model (GAM) and a Multi-Layer Perceptron (MLP). |
| title | Approximating Human Preferences Using a Multi-Judge Learned System |
| topic | Artificial Intelligence Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2510.25884 |