Evaluating the Prompt Steerability of Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Miehling, Erik, Desmond, Michael, Ramamurthy, Karthikeyan Natesan, Daly, Elizabeth M., Dognin, Pierre, Rios, Jesus, Bouneffouf, Djallel, Liu, Miao
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913692056027136
author Miehling, Erik
Desmond, Michael
Ramamurthy, Karthikeyan Natesan
Daly, Elizabeth M.
Dognin, Pierre
Rios, Jesus
Bouneffouf, Djallel
Liu, Miao
author_facet Miehling, Erik
Desmond, Michael
Ramamurthy, Karthikeyan Natesan
Daly, Elizabeth M.
Dognin, Pierre
Rios, Jesus
Bouneffouf, Djallel
Liu, Miao
contents Building pluralistic AI requires designing models that are able to be shaped to represent a wide range of value systems and cultures. Achieving this requires first being able to evaluate the degree to which a given model is capable of reflecting various personas. To this end, we propose a benchmark for evaluating the steerability of model personas as a function of prompting. Our design is based on a formal definition of prompt steerability, which analyzes the degree to which a model's joint behavioral distribution can be shifted from its baseline. By defining steerability indices and inspecting how these indices change as a function of steering effort, we can estimate the steerability of a model across various persona dimensions and directions. Our benchmark reveals that the steerability of many current models is limited -- due to both a skew in their baseline behavior and an asymmetry in their steerability across many persona dimensions. We release an implementation of our benchmark at https://github.com/IBM/prompt-steering.
format Preprint
id arxiv_https___arxiv_org_abs_2411_12405
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Evaluating the Prompt Steerability of Large Language Models
Miehling, Erik
Desmond, Michael
Ramamurthy, Karthikeyan Natesan
Daly, Elizabeth M.
Dognin, Pierre
Rios, Jesus
Bouneffouf, Djallel
Liu, Miao
Computation and Language
Artificial Intelligence
Human-Computer Interaction
Building pluralistic AI requires designing models that are able to be shaped to represent a wide range of value systems and cultures. Achieving this requires first being able to evaluate the degree to which a given model is capable of reflecting various personas. To this end, we propose a benchmark for evaluating the steerability of model personas as a function of prompting. Our design is based on a formal definition of prompt steerability, which analyzes the degree to which a model's joint behavioral distribution can be shifted from its baseline. By defining steerability indices and inspecting how these indices change as a function of steering effort, we can estimate the steerability of a model across various persona dimensions and directions. Our benchmark reveals that the steerability of many current models is limited -- due to both a skew in their baseline behavior and an asymmetry in their steerability across many persona dimensions. We release an implementation of our benchmark at https://github.com/IBM/prompt-steering.
title Evaluating the Prompt Steerability of Large Language Models
topic Computation and Language
Artificial Intelligence
Human-Computer Interaction
url https://arxiv.org/abs/2411.12405