Does Liking Yellow Imply Driving a School Bus? Semantic Leakage in Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gonen, Hila, Blevins, Terra, Liu, Alisa, Zettlemoyer, Luke, Smith, Noah A.
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912378034061312
author Gonen, Hila
Blevins, Terra
Liu, Alisa
Zettlemoyer, Luke
Smith, Noah A.
author_facet Gonen, Hila
Blevins, Terra
Liu, Alisa
Zettlemoyer, Luke
Smith, Noah A.
contents Despite their wide adoption, the biases and unintended behaviors of language models remain poorly understood. In this paper, we identify and characterize a phenomenon never discussed before, which we call semantic leakage, where models leak irrelevant information from the prompt into the generation in unexpected ways. We propose an evaluation setting to detect semantic leakage both by humans and automatically, curate a diverse test suite for diagnosing this behavior, and measure significant semantic leakage in 13 flagship models. We also show that models exhibit semantic leakage in languages besides English and across different settings and generation scenarios. This discovery highlights yet another type of bias in language models that affects their generation patterns and behavior.
format Preprint
id arxiv_https___arxiv_org_abs_2408_06518
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Does Liking Yellow Imply Driving a School Bus? Semantic Leakage in Language Models
Gonen, Hila
Blevins, Terra
Liu, Alisa
Zettlemoyer, Luke
Smith, Noah A.
Computation and Language
Despite their wide adoption, the biases and unintended behaviors of language models remain poorly understood. In this paper, we identify and characterize a phenomenon never discussed before, which we call semantic leakage, where models leak irrelevant information from the prompt into the generation in unexpected ways. We propose an evaluation setting to detect semantic leakage both by humans and automatically, curate a diverse test suite for diagnosing this behavior, and measure significant semantic leakage in 13 flagship models. We also show that models exhibit semantic leakage in languages besides English and across different settings and generation scenarios. This discovery highlights yet another type of bias in language models that affects their generation patterns and behavior.
title Does Liking Yellow Imply Driving a School Bus? Semantic Leakage in Language Models
topic Computation and Language
url https://arxiv.org/abs/2408.06518