Let's Measure Information Step-by-Step: AI-Based Evaluation Beyond Vibes

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Robertson, Zachary, Koyejo, Sanmi
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915967934660608
author Robertson, Zachary
Koyejo, Sanmi
author_facet Robertson, Zachary
Koyejo, Sanmi
contents We evaluate artificial intelligence (AI) systems without ground truth by exploiting a link between strategic gaming and information loss. Building on established information theory, we analyze which mechanisms resist adversarial manipulation. This motivates mutual evaluation, where the overseer is treated as a strategic player estimating mutual information by prompting, making truthful agent reporting an optimal strategy. We show that certain f-divergences, such as total variation distance (TVD), maintain polynomial guarantees under attack, building on an established exponential barrier for estimating mutual information (MI) in worst-case certification settings. Under adversarial attacks, TVD-MI maintains effectiveness (area under the curve 0.70--0.77) while other approaches can decay toward chance, demonstrating that prompting the same system for information relationships rather than quality judgments can improve robustness. The mechanisms decompose pairwise evaluations into reliable item-level detection scores without ground truth, addressing a key limitation of standard peer prediction. Pre-registration: https://osf.io/c7pum
format Preprint
id arxiv_https___arxiv_org_abs_2508_05469
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Let's Measure Information Step-by-Step: AI-Based Evaluation Beyond Vibes
Robertson, Zachary
Koyejo, Sanmi
Machine Learning
Information Theory
We evaluate artificial intelligence (AI) systems without ground truth by exploiting a link between strategic gaming and information loss. Building on established information theory, we analyze which mechanisms resist adversarial manipulation. This motivates mutual evaluation, where the overseer is treated as a strategic player estimating mutual information by prompting, making truthful agent reporting an optimal strategy. We show that certain f-divergences, such as total variation distance (TVD), maintain polynomial guarantees under attack, building on an established exponential barrier for estimating mutual information (MI) in worst-case certification settings. Under adversarial attacks, TVD-MI maintains effectiveness (area under the curve 0.70--0.77) while other approaches can decay toward chance, demonstrating that prompting the same system for information relationships rather than quality judgments can improve robustness. The mechanisms decompose pairwise evaluations into reliable item-level detection scores without ground truth, addressing a key limitation of standard peer prediction. Pre-registration: https://osf.io/c7pum
title Let's Measure Information Step-by-Step: AI-Based Evaluation Beyond Vibes
topic Machine Learning
Information Theory
url https://arxiv.org/abs/2508.05469