STAR: SocioTechnical Approach to Red Teaming Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Weidinger, Laura, Mellor, John, Pegueroles, Bernat Guillen, Marchal, Nahema, Kumar, Ravin, Lum, Kristian, Akbulut, Canfer, Diaz, Mark, Bergman, Stevie, Rodriguez, Mikel, Rieser, Verena, Isaac, William
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929554771148800
author Weidinger, Laura
Mellor, John
Pegueroles, Bernat Guillen
Marchal, Nahema
Kumar, Ravin
Lum, Kristian
Akbulut, Canfer
Diaz, Mark
Bergman, Stevie
Rodriguez, Mikel
Rieser, Verena
Isaac, William
author_facet Weidinger, Laura
Mellor, John
Pegueroles, Bernat Guillen
Marchal, Nahema
Kumar, Ravin
Lum, Kristian
Akbulut, Canfer
Diaz, Mark
Bergman, Stevie
Rodriguez, Mikel
Rieser, Verena
Isaac, William
contents This research introduces STAR, a sociotechnical framework that improves on current best practices for red teaming safety of large language models. STAR makes two key contributions: it enhances steerability by generating parameterised instructions for human red teamers, leading to improved coverage of the risk surface. Parameterised instructions also provide more detailed insights into model failures at no increased cost. Second, STAR improves signal quality by matching demographics to assess harms for specific groups, resulting in more sensitive annotations. STAR further employs a novel step of arbitration to leverage diverse viewpoints and improve label reliability, treating disagreement not as noise but as a valuable contribution to signal quality.
format Preprint
id arxiv_https___arxiv_org_abs_2406_11757
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle STAR: SocioTechnical Approach to Red Teaming Language Models
Weidinger, Laura
Mellor, John
Pegueroles, Bernat Guillen
Marchal, Nahema
Kumar, Ravin
Lum, Kristian
Akbulut, Canfer
Diaz, Mark
Bergman, Stevie
Rodriguez, Mikel
Rieser, Verena
Isaac, William
Artificial Intelligence
Computation and Language
Computers and Society
Human-Computer Interaction
This research introduces STAR, a sociotechnical framework that improves on current best practices for red teaming safety of large language models. STAR makes two key contributions: it enhances steerability by generating parameterised instructions for human red teamers, leading to improved coverage of the risk surface. Parameterised instructions also provide more detailed insights into model failures at no increased cost. Second, STAR improves signal quality by matching demographics to assess harms for specific groups, resulting in more sensitive annotations. STAR further employs a novel step of arbitration to leverage diverse viewpoints and improve label reliability, treating disagreement not as noise but as a valuable contribution to signal quality.
title STAR: SocioTechnical Approach to Red Teaming Language Models
topic Artificial Intelligence
Computation and Language
Computers and Society
Human-Computer Interaction
url https://arxiv.org/abs/2406.11757