PersonaTeaming: Exploring How Introducing Personas Can Improve Automated AI Red-Teaming
Fuente:
arXiv
Saved in:
| Main Authors: | Deng, Wesley Hanwen, Kim, Sunnie S. Y., Jha, Akshita, Holstein, Ken, Eslami, Motahhare, Wilcox, Lauren, Gatys, Leon A |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PersonaTeaming: Supporting Persona-Driven Red-Teaming for Generative AI
by: Deng, Wesley Hanwen, et al.
Published: (2026)
by: Deng, Wesley Hanwen, et al.
Published: (2026)
MIRAGE: Multi-model Interface for Reviewing and Auditing Generative Text-to-Image AI
by: Maldaner, Matheus Kunzler, et al.
Published: (2025)
by: Maldaner, Matheus Kunzler, et al.
Published: (2025)
Seeing Twice: How Side-by-Side T2I Comparison Changes Auditing Strategies
by: Maldaner, Matheus Kunzler, et al.
Published: (2025)
by: Maldaner, Matheus Kunzler, et al.
Published: (2025)
"I Don't Think RAI Applies to My Model'' -- Engaging Non-champions with Sticky Stories for Responsible AI Work
by: Nahar, Nadia, et al.
Published: (2025)
by: Nahar, Nadia, et al.
Published: (2025)
WeAudit: Scaffolding User Auditors and AI Practitioners in Auditing Generative AI
by: Deng, Wesley Hanwen, et al.
Published: (2025)
by: Deng, Wesley Hanwen, et al.
Published: (2025)
Red-Teaming for Generative AI: Silver Bullet or Security Theater?
by: Feffer, Michael, et al.
Published: (2024)
by: Feffer, Michael, et al.
Published: (2024)
Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning about Social Situations
by: Park, Eunkyu, et al.
Published: (2025)
by: Park, Eunkyu, et al.
Published: (2025)
Persona-Conditioned Adversarial Prompting: Multi-Identity Red-Teaming for Adversarial Discovery and Mitigation
by: Morasso, Cristian, et al.
Published: (2026)
by: Morasso, Cristian, et al.
Published: (2026)
Ruby Teaming: Improving Quality Diversity Search with Memory for Automated Red Teaming
by: Han, Vernon Toh Yan, et al.
Published: (2024)
by: Han, Vernon Toh Yan, et al.
Published: (2024)
Vipera: Towards systematic auditing of generative text-to-image models at scale
by: Huang, Yanwei, et al.
Published: (2025)
by: Huang, Yanwei, et al.
Published: (2025)
Persona-Conditioned Adversarial Prompting (PCAP): Multi-Identity Red-Teaming for Enhanced Adversarial Prompt Discovery
by: Morasso, Cristian, et al.
Published: (2026)
by: Morasso, Cristian, et al.
Published: (2026)
Automated Progressive Red Teaming
by: Jiang, Bojian, et al.
Published: (2024)
by: Jiang, Bojian, et al.
Published: (2024)
The Social Blindspot in Human-AI Collaboration: How Undetected AI Personas Reshape Team Dynamics
by: Yan, Lixiang, et al.
Published: (2025)
by: Yan, Lixiang, et al.
Published: (2025)
Critical or Compliant? The Double-Edged Sword of Reasoning in Chain-of-Thought Explanations
by: Park, Eunkyu, et al.
Published: (2025)
by: Park, Eunkyu, et al.
Published: (2025)
Prompt Optimization and Evaluation for LLM Automated Red Teaming
by: Freenor, Michael, et al.
Published: (2025)
by: Freenor, Michael, et al.
Published: (2025)
Red Teaming AI Red Teaming
by: Majumdar, Subhabrata, et al.
Published: (2025)
by: Majumdar, Subhabrata, et al.
Published: (2025)
MM-SCALE: Grounded Multimodal Moral Reasoning via Scalar Judgment and Listwise Alignment
by: Park, Eunkyu, et al.
Published: (2026)
by: Park, Eunkyu, et al.
Published: (2026)
Training a General Purpose Automated Red Teaming Model
by: Padmakumar, Aishwarya, et al.
Published: (2026)
by: Padmakumar, Aishwarya, et al.
Published: (2026)
The Automation Advantage in AI Red Teaming
by: Mulla, Rob, et al.
Published: (2025)
by: Mulla, Rob, et al.
Published: (2025)
Vipera: Blending Visual and LLM-Driven Guidance for Systematic Auditing of Text-to-Image Generative AI
by: Huang, Yanwei, et al.
Published: (2025)
by: Huang, Yanwei, et al.
Published: (2025)
Understanding Annotator Safety Policy with Interpretability
by: Oesterling, Alex, et al.
Published: (2026)
by: Oesterling, Alex, et al.
Published: (2026)
Exploring Straightforward Conversational Red-Teaming
by: Kour, George, et al.
Published: (2024)
by: Kour, George, et al.
Published: (2024)
Investigating Youth AI Auditing
by: Solyst, Jaemarie, et al.
Published: (2025)
by: Solyst, Jaemarie, et al.
Published: (2025)
Putting Privacy to the Test: Introducing Red Teaming for Research Data Anonymization
by: Jansen, Luisa, et al.
Published: (2026)
by: Jansen, Luisa, et al.
Published: (2026)
STAR-Teaming: A Strategy-Response Multiplex Network Approach to Automated LLM Red Teaming
by: Jung, MinJae, et al.
Published: (2026)
by: Jung, MinJae, et al.
Published: (2026)
Adaptive Instruction Composition for Automated LLM Red-Teaming
by: Zymet, Jesse, et al.
Published: (2026)
by: Zymet, Jesse, et al.
Published: (2026)
Anecdoctoring: Automated Red-Teaming Across Language and Place
by: Cuevas, Alejandro, et al.
Published: (2025)
by: Cuevas, Alejandro, et al.
Published: (2025)
How Can Teams Benefit From AI Team Members? Exploring the Effect of Generative AI on Decision‐Making Processes and Decision Quality in Team–AI Collaboration
by: Désirée Zercher, et al.
Published: (2025)
by: Désirée Zercher, et al.
Published: (2025)
RedCoder: Automated Multi-Turn Red Teaming for Code LLMs
by: Mo, Wenjie Jacky, et al.
Published: (2025)
by: Mo, Wenjie Jacky, et al.
Published: (2025)
Ferret: Faster and Effective Automated Red Teaming with Reward-Based Scoring Technique
by: Pala, Tej Deep, et al.
Published: (2024)
by: Pala, Tej Deep, et al.
Published: (2024)
Can LLMs Express Personality Across Cultures? Introducing CulturalPersonas for Evaluating Trait Alignment
by: Dey, Priyanka, et al.
Published: (2025)
by: Dey, Priyanka, et al.
Published: (2025)
Automated Red Teaming with GOAT: the Generative Offensive Agent Tester
by: Pavlova, Maya, et al.
Published: (2024)
by: Pavlova, Maya, et al.
Published: (2024)
Multi-lingual Multi-turn Automated Red Teaming for LLMs
by: Singhania, Abhishek, et al.
Published: (2025)
by: Singhania, Abhishek, et al.
Published: (2025)
Effective Automation to Support the Human Infrastructure in AI Red Teaming
by: Zhang, Alice Qian, et al.
Published: (2025)
by: Zhang, Alice Qian, et al.
Published: (2025)
MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring
by: Jotautaitė, Monika, et al.
Published: (2026)
by: Jotautaitė, Monika, et al.
Published: (2026)
Breaking Political Filter Bubbles via Social Comparison
by: Soliman, Nouran, et al.
Published: (2024)
by: Soliman, Nouran, et al.
Published: (2024)
Strategies for Designing Responsibly within a Capitalist Enterprise
by: Xie, Shixian, et al.
Published: (2026)
by: Xie, Shixian, et al.
Published: (2026)
BlueCodeAgent: A Blue Teaming Agent Enabled by Automated Red Teaming for CodeGen AI
by: Guo, Chengquan, et al.
Published: (2025)
by: Guo, Chengquan, et al.
Published: (2025)
SkillAttack: Automated Red Teaming of Agent Skills through Attack Path Refinement
by: Duan, Zenghao, et al.
Published: (2026)
by: Duan, Zenghao, et al.
Published: (2026)
Interactivity x Explainability: Toward Understanding How Interactivity Can Improve Computer Vision Explanations
by: Panigrahi, Indu, et al.
Published: (2025)
by: Panigrahi, Indu, et al.
Published: (2025)
Similar Items
-
PersonaTeaming: Supporting Persona-Driven Red-Teaming for Generative AI
by: Deng, Wesley Hanwen, et al.
Published: (2026) -
MIRAGE: Multi-model Interface for Reviewing and Auditing Generative Text-to-Image AI
by: Maldaner, Matheus Kunzler, et al.
Published: (2025) -
Seeing Twice: How Side-by-Side T2I Comparison Changes Auditing Strategies
by: Maldaner, Matheus Kunzler, et al.
Published: (2025) -
"I Don't Think RAI Applies to My Model'' -- Engaging Non-champions with Sticky Stories for Responsible AI Work
by: Nahar, Nadia, et al.
Published: (2025) -
WeAudit: Scaffolding User Auditors and AI Practitioners in Auditing Generative AI
by: Deng, Wesley Hanwen, et al.
Published: (2025)