Forbidden Facts: An Investigation of Competing Objectives in Llama-2

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Tony T., Wang, Miles, Hariharan, Kaivalya, Shavit, Nir
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913181456138240
author Wang, Tony T.
Wang, Miles
Hariharan, Kaivalya
Shavit, Nir
author_facet Wang, Tony T.
Wang, Miles
Hariharan, Kaivalya
Shavit, Nir
contents LLMs often face competing pressures (for example helpfulness vs. harmlessness). To understand how models resolve such conflicts, we study Llama-2-chat models on the forbidden fact task. Specifically, we instruct Llama-2 to truthfully complete a factual recall statement while forbidding it from saying the correct answer. This often makes the model give incorrect answers. We decompose Llama-2 into 1000+ components, and rank each one with respect to how useful it is for forbidding the correct answer. We find that in aggregate, around 35 components are enough to reliably implement the full suppression behavior. However, these components are fairly heterogeneous and many operate using faulty heuristics. We discover that one of these heuristics can be exploited via a manually designed adversarial attack which we call The California Attack. Our results highlight some roadblocks standing in the way of being able to successfully interpret advanced ML systems. Project website available at https://forbiddenfacts.github.io .
format Preprint
id arxiv_https___arxiv_org_abs_2312_08793
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Forbidden Facts: An Investigation of Competing Objectives in Llama-2
Wang, Tony T.
Wang, Miles
Hariharan, Kaivalya
Shavit, Nir
Machine Learning
Artificial Intelligence
Computation and Language
Cryptography and Security
LLMs often face competing pressures (for example helpfulness vs. harmlessness). To understand how models resolve such conflicts, we study Llama-2-chat models on the forbidden fact task. Specifically, we instruct Llama-2 to truthfully complete a factual recall statement while forbidding it from saying the correct answer. This often makes the model give incorrect answers. We decompose Llama-2 into 1000+ components, and rank each one with respect to how useful it is for forbidding the correct answer. We find that in aggregate, around 35 components are enough to reliably implement the full suppression behavior. However, these components are fairly heterogeneous and many operate using faulty heuristics. We discover that one of these heuristics can be exploited via a manually designed adversarial attack which we call The California Attack. Our results highlight some roadblocks standing in the way of being able to successfully interpret advanced ML systems. Project website available at https://forbiddenfacts.github.io .
title Forbidden Facts: An Investigation of Competing Objectives in Llama-2
topic Machine Learning
Artificial Intelligence
Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2312.08793