Risk and Response in Large Language Models: Evaluating Key Threat Categories

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Harandizadeh, Bahareh, Salinas, Abel, Morstatter, Fred
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910378545381376
author Harandizadeh, Bahareh
Salinas, Abel
Morstatter, Fred
author_facet Harandizadeh, Bahareh
Salinas, Abel
Morstatter, Fred
contents This paper explores the pressing issue of risk assessment in Large Language Models (LLMs) as they become increasingly prevalent in various applications. Focusing on how reward models, which are designed to fine-tune pretrained LLMs to align with human values, perceive and categorize different types of risks, we delve into the challenges posed by the subjective nature of preference-based training data. By utilizing the Anthropic Red-team dataset, we analyze major risk categories, including Information Hazards, Malicious Uses, and Discrimination/Hateful content. Our findings indicate that LLMs tend to consider Information Hazards less harmful, a finding confirmed by a specially developed regression model. Additionally, our analysis shows that LLMs respond less stringently to Information Hazards compared to other risks. The study further reveals a significant vulnerability of LLMs to jailbreaking attacks in Information Hazard scenarios, highlighting a critical security concern in LLM risk assessment and emphasizing the need for improved AI safety measures.
format Preprint
id arxiv_https___arxiv_org_abs_2403_14988
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Risk and Response in Large Language Models: Evaluating Key Threat Categories
Harandizadeh, Bahareh
Salinas, Abel
Morstatter, Fred
Computation and Language
This paper explores the pressing issue of risk assessment in Large Language Models (LLMs) as they become increasingly prevalent in various applications. Focusing on how reward models, which are designed to fine-tune pretrained LLMs to align with human values, perceive and categorize different types of risks, we delve into the challenges posed by the subjective nature of preference-based training data. By utilizing the Anthropic Red-team dataset, we analyze major risk categories, including Information Hazards, Malicious Uses, and Discrimination/Hateful content. Our findings indicate that LLMs tend to consider Information Hazards less harmful, a finding confirmed by a specially developed regression model. Additionally, our analysis shows that LLMs respond less stringently to Information Hazards compared to other risks. The study further reveals a significant vulnerability of LLMs to jailbreaking attacks in Information Hazard scenarios, highlighting a critical security concern in LLM risk assessment and emphasizing the need for improved AI safety measures.
title Risk and Response in Large Language Models: Evaluating Key Threat Categories
topic Computation and Language
url https://arxiv.org/abs/2403.14988