Branching Out: Broadening AI Measurement and Evaluation with Measurement Trees

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Greenberg, Craig, Hall, Patrick, Jensen, Theodore, Greene, Kristen, Amironesei, Razvan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912618713710592
author Greenberg, Craig
Hall, Patrick
Jensen, Theodore
Greene, Kristen
Amironesei, Razvan
author_facet Greenberg, Craig
Hall, Patrick
Jensen, Theodore
Greene, Kristen
Amironesei, Razvan
contents This paper introduces \textit{measurement trees}, a novel class of metrics designed to combine various constructs into an interpretable multi-level representation of a measurand. Unlike conventional metrics that yield single values, vectors, surfaces, or categories, measurement trees produce a hierarchical directed graph in which each node summarizes its children through user-defined aggregation methods. In response to recent calls to expand the scope of AI system evaluation, measurement trees enhance metric transparency and facilitate the integration of heterogeneous evidence, including, e.g., agentic, business, energy-efficiency, sociotechnical, or security signals. We present definitions and examples, demonstrate practical utility through a large-scale measurement exercise, and provide accompanying open-source Python code. By operationalizing a transparent approach to measurement of complex constructs, this work offers a principled foundation for broader and more interpretable AI evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2509_26632
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Branching Out: Broadening AI Measurement and Evaluation with Measurement Trees
Greenberg, Craig
Hall, Patrick
Jensen, Theodore
Greene, Kristen
Amironesei, Razvan
Artificial Intelligence
This paper introduces \textit{measurement trees}, a novel class of metrics designed to combine various constructs into an interpretable multi-level representation of a measurand. Unlike conventional metrics that yield single values, vectors, surfaces, or categories, measurement trees produce a hierarchical directed graph in which each node summarizes its children through user-defined aggregation methods. In response to recent calls to expand the scope of AI system evaluation, measurement trees enhance metric transparency and facilitate the integration of heterogeneous evidence, including, e.g., agentic, business, energy-efficiency, sociotechnical, or security signals. We present definitions and examples, demonstrate practical utility through a large-scale measurement exercise, and provide accompanying open-source Python code. By operationalizing a transparent approach to measurement of complex constructs, this work offers a principled foundation for broader and more interpretable AI evaluation.
title Branching Out: Broadening AI Measurement and Evaluation with Measurement Trees
topic Artificial Intelligence
url https://arxiv.org/abs/2509.26632