A Red Teaming Roadmap Towards System-Level Safety

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zifan, Knight, Christina Q., Kritz, Jeremy, Primack, Willow E., Michael, Julian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915332923326464
author Wang, Zifan
Knight, Christina Q.
Kritz, Jeremy
Primack, Willow E.
Michael, Julian
author_facet Wang, Zifan
Knight, Christina Q.
Kritz, Jeremy
Primack, Willow E.
Michael, Julian
contents Large Language Model (LLM) safeguards, which implement request refusals, have become a widely adopted mitigation strategy against misuse. At the intersection of adversarial machine learning and AI safety, safeguard red teaming has effectively identified critical vulnerabilities in state-of-the-art refusal-trained LLMs. However, in our view the many conference submissions on LLM red teaming do not, in aggregate, prioritize the right research problems. First, testing against clear product safety specifications should take a higher priority than abstract social biases or ethical principles. Second, red teaming should prioritize realistic threat models that represent the expanding risk landscape and what real attackers might do. Finally, we contend that system-level safety is a necessary step to move red teaming research forward, as AI models present new threats as well as affordances for threat mitigation (e.g., detection and banning of malicious users) once placed in a deployment context. Adopting these priorities will be necessary in order for red teaming research to adequately address the slate of new threats that rapid AI advances present today and will present in the very near future.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05376
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Red Teaming Roadmap Towards System-Level Safety
Wang, Zifan
Knight, Christina Q.
Kritz, Jeremy
Primack, Willow E.
Michael, Julian
Cryptography and Security
Artificial Intelligence
Large Language Model (LLM) safeguards, which implement request refusals, have become a widely adopted mitigation strategy against misuse. At the intersection of adversarial machine learning and AI safety, safeguard red teaming has effectively identified critical vulnerabilities in state-of-the-art refusal-trained LLMs. However, in our view the many conference submissions on LLM red teaming do not, in aggregate, prioritize the right research problems. First, testing against clear product safety specifications should take a higher priority than abstract social biases or ethical principles. Second, red teaming should prioritize realistic threat models that represent the expanding risk landscape and what real attackers might do. Finally, we contend that system-level safety is a necessary step to move red teaming research forward, as AI models present new threats as well as affordances for threat mitigation (e.g., detection and banning of malicious users) once placed in a deployment context. Adopting these priorities will be necessary in order for red teaming research to adequately address the slate of new threats that rapid AI advances present today and will present in the very near future.
title A Red Teaming Roadmap Towards System-Level Safety
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2506.05376