EigenBench: A Comparative Behavioral Measure of Value Alignment

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chang, Jonathn, Piff, Leonhard, Sana, Suvadip, Li, Jasmine X., Levine, Lionel
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908857754714112
author Chang, Jonathn
Piff, Leonhard
Sana, Suvadip
Li, Jasmine X.
Levine, Lionel
author_facet Chang, Jonathn
Piff, Leonhard
Sana, Suvadip
Li, Jasmine X.
Levine, Lionel
contents Aligning AI with human values is a pressing unsolved problem. To address the lack of quantitative metrics for value alignment, we propose EigenBench: a black-box method for comparatively benchmarking language models' values. Given an ensemble of models, a constitution describing a value system, and a dataset of scenarios, our method returns a vector of scores quantifying each model's alignment to the given constitution. To produce these scores, each model judges the outputs of other models across many scenarios, and these judgments are aggregated with EigenTrust (Kamvar et al., 2003), yielding scores that reflect a weighted consensus judgment of the whole ensemble. EigenBench uses no ground truth labels, as it is designed to quantify subjective traits for which reasonable judges may disagree on the correct label. Hence, to validate our method, we collect human judgments on the same ensemble of models and show that EigenBench's judgments align closely with those of human evaluators. We further demonstrate that EigenBench can recover model rankings on the GPQA benchmark without access to objective labels, supporting its viability as a framework for evaluating subjective values for which no ground truths exist. The code is available at https://github.com/jchang153/EigenBench.
format Preprint
id arxiv_https___arxiv_org_abs_2509_01938
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EigenBench: A Comparative Behavioral Measure of Value Alignment
Chang, Jonathn
Piff, Leonhard
Sana, Suvadip
Li, Jasmine X.
Levine, Lionel
Artificial Intelligence
Computation and Language
Computers and Society
Machine Learning
Aligning AI with human values is a pressing unsolved problem. To address the lack of quantitative metrics for value alignment, we propose EigenBench: a black-box method for comparatively benchmarking language models' values. Given an ensemble of models, a constitution describing a value system, and a dataset of scenarios, our method returns a vector of scores quantifying each model's alignment to the given constitution. To produce these scores, each model judges the outputs of other models across many scenarios, and these judgments are aggregated with EigenTrust (Kamvar et al., 2003), yielding scores that reflect a weighted consensus judgment of the whole ensemble. EigenBench uses no ground truth labels, as it is designed to quantify subjective traits for which reasonable judges may disagree on the correct label. Hence, to validate our method, we collect human judgments on the same ensemble of models and show that EigenBench's judgments align closely with those of human evaluators. We further demonstrate that EigenBench can recover model rankings on the GPQA benchmark without access to objective labels, supporting its viability as a framework for evaluating subjective values for which no ground truths exist. The code is available at https://github.com/jchang153/EigenBench.
title EigenBench: A Comparative Behavioral Measure of Value Alignment
topic Artificial Intelligence
Computation and Language
Computers and Society
Machine Learning
url https://arxiv.org/abs/2509.01938