PyRIT: A Framework for Security Risk Identification and Red Teaming in Generative AI System

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Munoz, Gary D. Lopez, Minnich, Amanda J., Lutz, Roman, Lundeen, Richard, Dheekonda, Raja Sekhar Rao, Chikanov, Nina, Jagdagdorj, Bolor-Erdene, Pouliot, Martin, Chawla, Shiven, Maxwell, Whitney, Bullwinkel, Blake, Pratt, Katherine, de Gruyter, Joris, Siska, Charlotte, Bryan, Pete, Westerhoff, Tori, Kawaguchi, Chang, Seifert, Christian, Kumar, Ram Shankar Siva, Zunger, Yonatan
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929526669312000
author Munoz, Gary D. Lopez
Minnich, Amanda J.
Lutz, Roman
Lundeen, Richard
Dheekonda, Raja Sekhar Rao
Chikanov, Nina
Jagdagdorj, Bolor-Erdene
Pouliot, Martin
Chawla, Shiven
Maxwell, Whitney
Bullwinkel, Blake
Pratt, Katherine
de Gruyter, Joris
Siska, Charlotte
Bryan, Pete
Westerhoff, Tori
Kawaguchi, Chang
Seifert, Christian
Kumar, Ram Shankar Siva
Zunger, Yonatan
author_facet Munoz, Gary D. Lopez
Minnich, Amanda J.
Lutz, Roman
Lundeen, Richard
Dheekonda, Raja Sekhar Rao
Chikanov, Nina
Jagdagdorj, Bolor-Erdene
Pouliot, Martin
Chawla, Shiven
Maxwell, Whitney
Bullwinkel, Blake
Pratt, Katherine
de Gruyter, Joris
Siska, Charlotte
Bryan, Pete
Westerhoff, Tori
Kawaguchi, Chang
Seifert, Christian
Kumar, Ram Shankar Siva
Zunger, Yonatan
contents Generative Artificial Intelligence (GenAI) is becoming ubiquitous in our daily lives. The increase in computational power and data availability has led to a proliferation of both single- and multi-modal models. As the GenAI ecosystem matures, the need for extensible and model-agnostic risk identification frameworks is growing. To meet this need, we introduce the Python Risk Identification Toolkit (PyRIT), an open-source framework designed to enhance red teaming efforts in GenAI systems. PyRIT is a model- and platform-agnostic tool that enables red teamers to probe for and identify novel harms, risks, and jailbreaks in multimodal generative AI models. Its composable architecture facilitates the reuse of core building blocks and allows for extensibility to future models and modalities. This paper details the challenges specific to red teaming generative AI systems, the development and features of PyRIT, and its practical applications in real-world scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2410_02828
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PyRIT: A Framework for Security Risk Identification and Red Teaming in Generative AI System
Munoz, Gary D. Lopez
Minnich, Amanda J.
Lutz, Roman
Lundeen, Richard
Dheekonda, Raja Sekhar Rao
Chikanov, Nina
Jagdagdorj, Bolor-Erdene
Pouliot, Martin
Chawla, Shiven
Maxwell, Whitney
Bullwinkel, Blake
Pratt, Katherine
de Gruyter, Joris
Siska, Charlotte
Bryan, Pete
Westerhoff, Tori
Kawaguchi, Chang
Seifert, Christian
Kumar, Ram Shankar Siva
Zunger, Yonatan
Cryptography and Security
Artificial Intelligence
Computation and Language
Generative Artificial Intelligence (GenAI) is becoming ubiquitous in our daily lives. The increase in computational power and data availability has led to a proliferation of both single- and multi-modal models. As the GenAI ecosystem matures, the need for extensible and model-agnostic risk identification frameworks is growing. To meet this need, we introduce the Python Risk Identification Toolkit (PyRIT), an open-source framework designed to enhance red teaming efforts in GenAI systems. PyRIT is a model- and platform-agnostic tool that enables red teamers to probe for and identify novel harms, risks, and jailbreaks in multimodal generative AI models. Its composable architecture facilitates the reuse of core building blocks and allows for extensibility to future models and modalities. This paper details the challenges specific to red teaming generative AI systems, the development and features of PyRIT, and its practical applications in real-world scenarios.
title PyRIT: A Framework for Security Risk Identification and Red Teaming in Generative AI System
topic Cryptography and Security
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2410.02828