Policy Maps: Tools for Guiding the Unbounded Space of LLM Behaviors

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lam, Michelle S., Hohman, Fred, Moritz, Dominik, Bigham, Jeffrey P., Holstein, Kenneth, Kery, Mary Beth
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908474203439104
author Lam, Michelle S.
Hohman, Fred
Moritz, Dominik
Bigham, Jeffrey P.
Holstein, Kenneth
Kery, Mary Beth
author_facet Lam, Michelle S.
Hohman, Fred
Moritz, Dominik
Bigham, Jeffrey P.
Holstein, Kenneth
Kery, Mary Beth
contents AI policy sets boundaries on acceptable behavior for AI models, but this is challenging in the context of large language models (LLMs): how do you ensure coverage over a vast behavior space? We introduce policy maps, an approach to AI policy design inspired by the practice of physical mapmaking. Instead of aiming for full coverage, policy maps aid effective navigation through intentional design choices about which aspects to capture and which to abstract away. With Policy Projector, an interactive tool for designing LLM policy maps, an AI practitioner can survey the landscape of model input-output pairs, define custom regions (e.g., "violence"), and navigate these regions with if-then policy rules that can act on LLM outputs (e.g., if output contains "violence" and "graphic details," then rewrite without "graphic details"). Policy Projector supports interactive policy authoring using LLM classification and steering and a map visualization reflecting the AI practitioner's work. In an evaluation with 12 AI safety experts, our system helps policy designers craft policies around problematic model behaviors such as incorrect gender assumptions and handling of immediate physical safety threats.
format Preprint
id arxiv_https___arxiv_org_abs_2409_18203
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Policy Maps: Tools for Guiding the Unbounded Space of LLM Behaviors
Lam, Michelle S.
Hohman, Fred
Moritz, Dominik
Bigham, Jeffrey P.
Holstein, Kenneth
Kery, Mary Beth
Human-Computer Interaction
Artificial Intelligence
Computation and Language
Machine Learning
AI policy sets boundaries on acceptable behavior for AI models, but this is challenging in the context of large language models (LLMs): how do you ensure coverage over a vast behavior space? We introduce policy maps, an approach to AI policy design inspired by the practice of physical mapmaking. Instead of aiming for full coverage, policy maps aid effective navigation through intentional design choices about which aspects to capture and which to abstract away. With Policy Projector, an interactive tool for designing LLM policy maps, an AI practitioner can survey the landscape of model input-output pairs, define custom regions (e.g., "violence"), and navigate these regions with if-then policy rules that can act on LLM outputs (e.g., if output contains "violence" and "graphic details," then rewrite without "graphic details"). Policy Projector supports interactive policy authoring using LLM classification and steering and a map visualization reflecting the AI practitioner's work. In an evaluation with 12 AI safety experts, our system helps policy designers craft policies around problematic model behaviors such as incorrect gender assumptions and handling of immediate physical safety threats.
title Policy Maps: Tools for Guiding the Unbounded Space of LLM Behaviors
topic Human-Computer Interaction
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2409.18203