Automated Red Teaming with GOAT: the Generative Offensive Agent Tester
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Pavlova, Maya, Brinkman, Erik, Iyer, Krithika, Albiero, Vitor, Bitton, Joanna, Nguyen, Hailey, Li, Joe, Ferrer, Cristian Canton, Evtimov, Ivan, Grattafiori, Aaron |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
FERRET: Framework for Expansion Reliant Red Teaming
par: Mehrabi, Ninareh, et autres
Publié: (2026)
par: Mehrabi, Ninareh, et autres
Publié: (2026)
WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
par: Evtimov, Ivan, et autres
Publié: (2025)
par: Evtimov, Ivan, et autres
Publié: (2025)
Towards Red Teaming in Multimodal and Multilingual Translation
par: Ropers, Christophe, et autres
Publié: (2024)
par: Ropers, Christophe, et autres
Publié: (2024)
Gradient-based Jailbreak Images for Multimodal Fusion Models
par: Rando, Javier, et autres
Publié: (2024)
par: Rando, Javier, et autres
Publié: (2024)
AgentDAM: Privacy Leakage Evaluation for Autonomous Web Agents
par: Zharmagambetov, Arman, et autres
Publié: (2025)
par: Zharmagambetov, Arman, et autres
Publié: (2025)
Probabilistic 3D Correspondence Prediction from Sparse Unsegmented Images
par: Iyer, Krithika, et autres
Publié: (2024)
par: Iyer, Krithika, et autres
Publié: (2024)
Automated Progressive Red Teaming
par: Jiang, Bojian, et autres
Publié: (2024)
par: Jiang, Bojian, et autres
Publié: (2024)
Code Llama: Open Foundation Models for Code
par: Rozière, Baptiste, et autres
Publié: (2023)
par: Rozière, Baptiste, et autres
Publié: (2023)
Weakly Supervised Bayesian Shape Modeling from Unsegmented Medical Images
par: Adams, Jadie, et autres
Publié: (2024)
par: Adams, Jadie, et autres
Publié: (2024)
LEDA: Log-Euclidean Diffeomorphism Autoencoder for Efficient Statistical Analysis of Diffeomorphisms
par: Iyer, Krithika, et autres
Publié: (2024)
par: Iyer, Krithika, et autres
Publié: (2024)
The Automation Advantage in AI Red Teaming
par: Mulla, Rob, et autres
Publié: (2025)
par: Mulla, Rob, et autres
Publié: (2025)
Can student loans improve accessibility to higher education and student performance? : n impact study of the case of SOFES, Mexico / Erik Canton, Andreas Blom
par: Canton, Erik
Publié: (2004)
par: Canton, Erik
Publié: (2004)
Calidad ósea, el rol de los diferentes métodos complementarios
par: Alejandro Albiero
Publié: (2020)
par: Alejandro Albiero
Publié: (2020)
BreachSeek: A Multi-Agent Automated Penetration Tester
par: Alshehri, Ibrahim, et autres
Publié: (2024)
par: Alshehri, Ibrahim, et autres
Publié: (2024)
SCorP: Statistics-Informed Dense Correspondence Prediction Directly from Unsegmented Medical Images
par: Iyer, Krithika, et autres
Publié: (2024)
par: Iyer, Krithika, et autres
Publié: (2024)
RedCoder: Automated Multi-Turn Red Teaming for Code LLMs
par: Mo, Wenjie Jacky, et autres
Publié: (2025)
par: Mo, Wenjie Jacky, et autres
Publié: (2025)
Ruby Teaming: Improving Quality Diversity Search with Memory for Automated Red Teaming
par: Han, Vernon Toh Yan, et autres
Publié: (2024)
par: Han, Vernon Toh Yan, et autres
Publié: (2024)
Red Teaming AI Red Teaming
par: Majumdar, Subhabrata, et autres
Publié: (2025)
par: Majumdar, Subhabrata, et autres
Publié: (2025)
Adaptive Instruction Composition for Automated LLM Red-Teaming
par: Zymet, Jesse, et autres
Publié: (2026)
par: Zymet, Jesse, et autres
Publié: (2026)
Anecdoctoring: Automated Red-Teaming Across Language and Place
par: Cuevas, Alejandro, et autres
Publié: (2025)
par: Cuevas, Alejandro, et autres
Publié: (2025)
Prompt Optimization and Evaluation for LLM Automated Red Teaming
par: Freenor, Michael, et autres
Publié: (2025)
par: Freenor, Michael, et autres
Publié: (2025)
Média dos valores da frase em crianças com desvio fonológico evolutivo
par: Jamile Konzen Albiero
Publié: (2011)
par: Jamile Konzen Albiero
Publié: (2011)
Mesh2SSM++: A Probabilistic Framework for Unsupervised Learning of Statistical Shape Model of Anatomies from Surface Meshes
par: Iyer, Krithika, et autres
Publié: (2025)
par: Iyer, Krithika, et autres
Publié: (2025)
From Automation to Augmentation: A Framework for Designing Human-Centric Work Environments in Society 5.0
par: Maya, Cristian Espinal
Publié: (2026)
par: Maya, Cristian Espinal
Publié: (2026)
RAN Tester UE: An Automated Declarative UE Centric Security Testing Platform
par: Ueltschey, Charles Marion, et autres
Publié: (2025)
par: Ueltschey, Charles Marion, et autres
Publié: (2025)
Design and Development of an Automated Contact Angle Tester (ACAT) for Surface Wettability Measurement
par: Burgess, Connor, et autres
Publié: (2025)
par: Burgess, Connor, et autres
Publié: (2025)
Abstractive Red-Teaming of Language Model Character
par: Rahn, Nate, et autres
Publié: (2026)
par: Rahn, Nate, et autres
Publié: (2026)
Multi-lingual Multi-turn Automated Red Teaming for LLMs
par: Singhania, Abhishek, et autres
Publié: (2025)
par: Singhania, Abhishek, et autres
Publié: (2025)
Effective Automation to Support the Human Infrastructure in AI Red Teaming
par: Zhang, Alice Qian, et autres
Publié: (2025)
par: Zhang, Alice Qian, et autres
Publié: (2025)
Training a General Purpose Automated Red Teaming Model
par: Padmakumar, Aishwarya, et autres
Publié: (2026)
par: Padmakumar, Aishwarya, et autres
Publié: (2026)
MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring
par: Jotautaitė, Monika, et autres
Publié: (2026)
par: Jotautaitė, Monika, et autres
Publié: (2026)
GOAT: A Training Framework for Goal-Oriented Agent with Tools
par: Min, Hyunji, et autres
Publié: (2025)
par: Min, Hyunji, et autres
Publié: (2025)
GOAT-Bench: A Benchmark for Multi-Modal Lifelong Navigation
par: Khanna, Mukul, et autres
Publié: (2024)
par: Khanna, Mukul, et autres
Publié: (2024)
GOAT: A Global Optimization Algorithm for Molecules and Atomic Clusters
par: Bernardo de Souza
Publié: (2025)
par: Bernardo de Souza
Publié: (2025)
GOAT: A Global Optimization Algorithm for Molecules and Atomic Clusters
par: Bernardo de Souza
Publié: (2025)
par: Bernardo de Souza
Publié: (2025)
PersonaTeaming: Exploring How Introducing Personas Can Improve Automated AI Red-Teaming
par: Deng, Wesley Hanwen, et autres
Publié: (2025)
par: Deng, Wesley Hanwen, et autres
Publié: (2025)
STAR-Teaming: A Strategy-Response Multiplex Network Approach to Automated LLM Red Teaming
par: Jung, MinJae, et autres
Publié: (2026)
par: Jung, MinJae, et autres
Publié: (2026)
On the Structure of Replicable Hypothesis Testers
par: Aamand, Anders, et autres
Publié: (2025)
par: Aamand, Anders, et autres
Publié: (2025)
Demo: ViolentUTF as An Accessible Platform for Generative AI Red Teaming
par: Nguyen, Tam n.
Publié: (2025)
par: Nguyen, Tam n.
Publié: (2025)
Detecting Stylistic Fingerprints of Large Language Models
par: Bitton, Yehonatan, et autres
Publié: (2025)
par: Bitton, Yehonatan, et autres
Publié: (2025)
Documents similaires
-
FERRET: Framework for Expansion Reliant Red Teaming
par: Mehrabi, Ninareh, et autres
Publié: (2026) -
WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
par: Evtimov, Ivan, et autres
Publié: (2025) -
Towards Red Teaming in Multimodal and Multilingual Translation
par: Ropers, Christophe, et autres
Publié: (2024) -
Gradient-based Jailbreak Images for Multimodal Fusion Models
par: Rando, Javier, et autres
Publié: (2024) -
AgentDAM: Privacy Leakage Evaluation for Autonomous Web Agents
par: Zharmagambetov, Arman, et autres
Publié: (2025)