Approximating Human Preferences Using a Multi-Judge Learned System

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sprejer, Eitán, Avalos, Fernando, Bernardi, Augusto, Faustino, Jose Pedro Brito de Azevedo, Haimes, Jacob, Oozeer, Narmeen Fatimah
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908619789828096
author Sprejer, Eitán
Avalos, Fernando
Bernardi, Augusto
Faustino, Jose Pedro Brito de Azevedo
Haimes, Jacob
Oozeer, Narmeen Fatimah
author_facet Sprejer, Eitán
Avalos, Fernando
Bernardi, Augusto
Faustino, Jose Pedro Brito de Azevedo
Haimes, Jacob
Oozeer, Narmeen Fatimah
contents Aligning LLM-based judges with human preferences is a significant challenge, as they are difficult to calibrate and often suffer from rubric sensitivity, bias, and instability. Overcoming this challenge advances key applications, such as creating reliable reward models for Reinforcement Learning from Human Feedback (RLHF) and building effective routing systems that select the best-suited model for a given user query. In this work, we propose a framework for modeling diverse, persona-based preferences by learning to aggregate outputs from multiple rubric-conditioned judges. We investigate the performance of this approach against naive baselines and assess its robustness through case studies on both human and LLM-judges biases. Our primary contributions include a persona-based method for synthesizing preference labels at scale and two distinct implementations of our aggregator: Generalized Additive Model (GAM) and a Multi-Layer Perceptron (MLP).
format Preprint
id arxiv_https___arxiv_org_abs_2510_25884
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Approximating Human Preferences Using a Multi-Judge Learned System
Sprejer, Eitán
Avalos, Fernando
Bernardi, Augusto
Faustino, Jose Pedro Brito de Azevedo
Haimes, Jacob
Oozeer, Narmeen Fatimah
Artificial Intelligence
Computation and Language
Machine Learning
Aligning LLM-based judges with human preferences is a significant challenge, as they are difficult to calibrate and often suffer from rubric sensitivity, bias, and instability. Overcoming this challenge advances key applications, such as creating reliable reward models for Reinforcement Learning from Human Feedback (RLHF) and building effective routing systems that select the best-suited model for a given user query. In this work, we propose a framework for modeling diverse, persona-based preferences by learning to aggregate outputs from multiple rubric-conditioned judges. We investigate the performance of this approach against naive baselines and assess its robustness through case studies on both human and LLM-judges biases. Our primary contributions include a persona-based method for synthesizing preference labels at scale and two distinct implementations of our aggregator: Generalized Additive Model (GAM) and a Multi-Layer Perceptron (MLP).
title Approximating Human Preferences Using a Multi-Judge Learned System
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2510.25884