Dimensions of Generative AI Evaluation Design

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dow, P. Alex, Vaughan, Jennifer Wortman, Barocas, Solon, Atalla, Chad, Chouldechova, Alexandra, Wallach, Hanna
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912126772183040
author Dow, P. Alex
Vaughan, Jennifer Wortman
Barocas, Solon
Atalla, Chad
Chouldechova, Alexandra
Wallach, Hanna
author_facet Dow, P. Alex
Vaughan, Jennifer Wortman
Barocas, Solon
Atalla, Chad
Chouldechova, Alexandra
Wallach, Hanna
contents There are few principles or guidelines to ensure evaluations of generative AI (GenAI) models and systems are effective. To help address this gap, we propose a set of general dimensions that capture critical choices involved in GenAI evaluation design. These dimensions include the evaluation setting, the task type, the input source, the interaction style, the duration, the metric type, and the scoring method. By situating GenAI evaluations within these dimensions, we aim to guide decision-making during GenAI evaluation design and provide a structure for comparing different evaluations. We illustrate the utility of the proposed set of general dimensions using two examples: a hypothetical evaluation of the fairness of a GenAI system and three real-world GenAI evaluations of biological threats.
format Preprint
id arxiv_https___arxiv_org_abs_2411_12709
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Dimensions of Generative AI Evaluation Design
Dow, P. Alex
Vaughan, Jennifer Wortman
Barocas, Solon
Atalla, Chad
Chouldechova, Alexandra
Wallach, Hanna
Computers and Society
There are few principles or guidelines to ensure evaluations of generative AI (GenAI) models and systems are effective. To help address this gap, we propose a set of general dimensions that capture critical choices involved in GenAI evaluation design. These dimensions include the evaluation setting, the task type, the input source, the interaction style, the duration, the metric type, and the scoring method. By situating GenAI evaluations within these dimensions, we aim to guide decision-making during GenAI evaluation design and provide a structure for comparing different evaluations. We illustrate the utility of the proposed set of general dimensions using two examples: a hypothetical evaluation of the fairness of a GenAI system and three real-world GenAI evaluations of biological threats.
title Dimensions of Generative AI Evaluation Design
topic Computers and Society
url https://arxiv.org/abs/2411.12709