Logical Consistency Between Disagreeing Experts and Its Role in AI Safety

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Corrada-Emmanuel, Andrés
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914070332964864
author Corrada-Emmanuel, Andrés
author_facet Corrada-Emmanuel, Andrés
contents If two experts disagree on a test, we may conclude both cannot be 100 per cent correct. But if they completely agree, no possible evaluation can be excluded. This asymmetry in the utility of agreements versus disagreements is explored here by formalizing a logic of unsupervised evaluation for classifiers. Its core problem is computing the set of group evaluations that are logically consistent with how we observe them agreeing and disagreeing in their decisions. Statistical summaries of their aligned decisions are inputs into a Linear Programming problem in the integer space of possible correct or incorrect responses given true labels. Obvious logical constraints, such as, the number of correct responses cannot exceed the number of observed responses, are inequalities. But in addition, there are axioms, universally applicable linear equalities that apply to all finite tests. The practical and immediate utility of this approach to unsupervised evaluation using only logical consistency is demonstrated by building no-knowledge alarms that can detect when one or more LLMs-as-Judges are violating a minimum grading threshold specified by the user.
format Preprint
id arxiv_https___arxiv_org_abs_2510_00821
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Logical Consistency Between Disagreeing Experts and Its Role in AI Safety
Corrada-Emmanuel, Andrés
Artificial Intelligence
90C05, 68T27
I.2.3; F.4.1
If two experts disagree on a test, we may conclude both cannot be 100 per cent correct. But if they completely agree, no possible evaluation can be excluded. This asymmetry in the utility of agreements versus disagreements is explored here by formalizing a logic of unsupervised evaluation for classifiers. Its core problem is computing the set of group evaluations that are logically consistent with how we observe them agreeing and disagreeing in their decisions. Statistical summaries of their aligned decisions are inputs into a Linear Programming problem in the integer space of possible correct or incorrect responses given true labels. Obvious logical constraints, such as, the number of correct responses cannot exceed the number of observed responses, are inequalities. But in addition, there are axioms, universally applicable linear equalities that apply to all finite tests. The practical and immediate utility of this approach to unsupervised evaluation using only logical consistency is demonstrated by building no-knowledge alarms that can detect when one or more LLMs-as-Judges are violating a minimum grading threshold specified by the user.
title Logical Consistency Between Disagreeing Experts and Its Role in AI Safety
topic Artificial Intelligence
90C05, 68T27
I.2.3; F.4.1
url https://arxiv.org/abs/2510.00821