Contextualized Evaluations: Judging Language Model Responses to Underspecified Queries

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Malaviya, Chaitanya, Chang, Joseph Chee, Roth, Dan, Iyyer, Mohit, Yatskar, Mark, Lo, Kyle
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909621659107328
author Malaviya, Chaitanya
Chang, Joseph Chee
Roth, Dan
Iyyer, Mohit
Yatskar, Mark
Lo, Kyle
author_facet Malaviya, Chaitanya
Chang, Joseph Chee
Roth, Dan
Iyyer, Mohit
Yatskar, Mark
Lo, Kyle
contents Language model users often issue queries that lack specification, where the context under which a query was issued -- such as the user's identity, the query's intent, and the criteria for a response to be useful -- is not explicit. For instance, a good response to a subjective query like "What book should I read next?" would depend on the user's preferences, and a good response to an open-ended query like "How do antibiotics work against bacteria?" would depend on the user's expertise. This makes evaluation of responses to such queries an ill-posed task, as evaluators may make arbitrary judgments about the response quality. To remedy this, we present contextualized evaluations, a protocol that synthetically constructs context surrounding an underspecified query and provides it during evaluation. We find that the presence of context can 1) alter conclusions drawn from evaluation, even flipping benchmark rankings between model pairs, 2) nudge evaluators to make fewer judgments based on surface-level criteria, like style, and 3) provide new insights about model behavior across diverse contexts. Specifically, our procedure suggests a potential bias towards WEIRD (Western, Educated, Industrialized, Rich and Democratic) contexts in models' "default" responses and we find that models are not equally sensitive to following different contexts, even when they are provided in prompts.
format Preprint
id arxiv_https___arxiv_org_abs_2411_07237
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Contextualized Evaluations: Judging Language Model Responses to Underspecified Queries
Malaviya, Chaitanya
Chang, Joseph Chee
Roth, Dan
Iyyer, Mohit
Yatskar, Mark
Lo, Kyle
Computation and Language
Language model users often issue queries that lack specification, where the context under which a query was issued -- such as the user's identity, the query's intent, and the criteria for a response to be useful -- is not explicit. For instance, a good response to a subjective query like "What book should I read next?" would depend on the user's preferences, and a good response to an open-ended query like "How do antibiotics work against bacteria?" would depend on the user's expertise. This makes evaluation of responses to such queries an ill-posed task, as evaluators may make arbitrary judgments about the response quality. To remedy this, we present contextualized evaluations, a protocol that synthetically constructs context surrounding an underspecified query and provides it during evaluation. We find that the presence of context can 1) alter conclusions drawn from evaluation, even flipping benchmark rankings between model pairs, 2) nudge evaluators to make fewer judgments based on surface-level criteria, like style, and 3) provide new insights about model behavior across diverse contexts. Specifically, our procedure suggests a potential bias towards WEIRD (Western, Educated, Industrialized, Rich and Democratic) contexts in models' "default" responses and we find that models are not equally sensitive to following different contexts, even when they are provided in prompts.
title Contextualized Evaluations: Judging Language Model Responses to Underspecified Queries
topic Computation and Language
url https://arxiv.org/abs/2411.07237