Red-Teaming for Generative AI: Silver Bullet or Security Theater?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Feffer, Michael, Sinha, Anusha, Deng, Wesley Hanwen, Lipton, Zachary C., Heidari, Hoda |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PersonaTeaming: Supporting Persona-Driven Red-Teaming for Generative AI
von: Deng, Wesley Hanwen, et al.
Veröffentlicht: (2026)
von: Deng, Wesley Hanwen, et al.
Veröffentlicht: (2026)
The Generative AI Ethics Playbook
von: Smith, Jessie J., et al.
Veröffentlicht: (2024)
von: Smith, Jessie J., et al.
Veröffentlicht: (2024)
On the Pros and Cons of Active Learning for Moral Preference Elicitation
von: Keswani, Vijay, et al.
Veröffentlicht: (2024)
von: Keswani, Vijay, et al.
Veröffentlicht: (2024)
Not a Silver Bullet for Loneliness: How Attachment and Age Shape Intimacy with AI Companions
von: Ciriello, Raffaele, et al.
Veröffentlicht: (2026)
von: Ciriello, Raffaele, et al.
Veröffentlicht: (2026)
Effective Automation to Support the Human Infrastructure in AI Red Teaming
von: Zhang, Alice Qian, et al.
Veröffentlicht: (2025)
von: Zhang, Alice Qian, et al.
Veröffentlicht: (2025)
Investigating Youth AI Auditing
von: Solyst, Jaemarie, et al.
Veröffentlicht: (2025)
von: Solyst, Jaemarie, et al.
Veröffentlicht: (2025)
PersonaTeaming: Exploring How Introducing Personas Can Improve Automated AI Red-Teaming
von: Deng, Wesley Hanwen, et al.
Veröffentlicht: (2025)
von: Deng, Wesley Hanwen, et al.
Veröffentlicht: (2025)
BoilerTAI: A Platform for Enhancing Instruction Using Generative AI in Educational Forums
von: Sinha, Anvit, et al.
Veröffentlicht: (2024)
von: Sinha, Anvit, et al.
Veröffentlicht: (2024)
The "Who", "What", and "How" of Responsible AI Governance: A Systematic Review and Meta-Analysis of (Actor, Stage)-Specific Tools
von: Kuehnert, Blaine, et al.
Veröffentlicht: (2025)
von: Kuehnert, Blaine, et al.
Veröffentlicht: (2025)
OpenAI's Approach to External Red Teaming for AI Models and Systems
von: Ahmad, Lama, et al.
Veröffentlicht: (2025)
von: Ahmad, Lama, et al.
Veröffentlicht: (2025)
Legacy Procurement Practices Shape How U.S. Cities Govern AI: Understanding Government Employees' Practices, Challenges, and Needs
von: Johnson, Nari, et al.
Veröffentlicht: (2024)
von: Johnson, Nari, et al.
Veröffentlicht: (2024)
Can AI Model the Complexities of Human Moral Decision-Making? A Qualitative Study of Kidney Allocation Decisions
von: Keswani, Vijay, et al.
Veröffentlicht: (2025)
von: Keswani, Vijay, et al.
Veröffentlicht: (2025)
The Human Factor in AI Red Teaming: Perspectives from Social and Collaborative Computing
von: Zhang, Alice Qian, et al.
Veröffentlicht: (2024)
von: Zhang, Alice Qian, et al.
Veröffentlicht: (2024)
Moral Change or Noise? On Problems of Aligning AI With Temporally Unstable Human Feedback
von: Keswani, Vijay, et al.
Veröffentlicht: (2025)
von: Keswani, Vijay, et al.
Veröffentlicht: (2025)
Analyzing Security and Privacy Challenges in Generative AI Usage Guidelines for Higher Education
von: Ng, Bei Yi, et al.
Veröffentlicht: (2025)
von: Ng, Bei Yi, et al.
Veröffentlicht: (2025)
How AI Impacts Skill Formation
von: Shen, Judy Hanwen, et al.
Veröffentlicht: (2026)
von: Shen, Judy Hanwen, et al.
Veröffentlicht: (2026)
Troubling Taxonomies in GenAI Evaluation
von: Berman, Glen, et al.
Veröffentlicht: (2024)
von: Berman, Glen, et al.
Veröffentlicht: (2024)
Why Do Decision Makers (Not) Use AI? A Cross-Domain Analysis of Factors Impacting AI Adoption
von: Yu, Rebecca, et al.
Veröffentlicht: (2025)
von: Yu, Rebecca, et al.
Veröffentlicht: (2025)
MIRAGE: Multi-model Interface for Reviewing and Auditing Generative Text-to-Image AI
von: Maldaner, Matheus Kunzler, et al.
Veröffentlicht: (2025)
von: Maldaner, Matheus Kunzler, et al.
Veröffentlicht: (2025)
Disclosure or Marketing? Analyzing the Efficacy of Vendor Self-reports for Vetting Public-sector AI
von: Kuehnert, Blaine, et al.
Veröffentlicht: (2026)
von: Kuehnert, Blaine, et al.
Veröffentlicht: (2026)
"Till I can get my satisfaction": Open Questions in the Public Desire to Punish AI
von: Ungless, Eddie L., et al.
Veröffentlicht: (2025)
von: Ungless, Eddie L., et al.
Veröffentlicht: (2025)
The RIGID Framework: Research-Integrated, Generative AI-Mediated Instructional Design
von: Kwak, Yerin, et al.
Veröffentlicht: (2026)
von: Kwak, Yerin, et al.
Veröffentlicht: (2026)
On The Stability of Moral Preferences: A Problem with Computational Elicitation Methods
von: Boerstler, Kyle, et al.
Veröffentlicht: (2024)
von: Boerstler, Kyle, et al.
Veröffentlicht: (2024)
ChatISA: A Prompt-Engineered, In-House Multi-Modal Generative AI Chatbot for Information Systems Education
von: Megahed, Fadel M., et al.
Veröffentlicht: (2024)
von: Megahed, Fadel M., et al.
Veröffentlicht: (2024)
WeAudit: Scaffolding User Auditors and AI Practitioners in Auditing Generative AI
von: Deng, Wesley Hanwen, et al.
Veröffentlicht: (2025)
von: Deng, Wesley Hanwen, et al.
Veröffentlicht: (2025)
Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South
von: Rastogi, Charvi, et al.
Veröffentlicht: (2026)
von: Rastogi, Charvi, et al.
Veröffentlicht: (2026)
Nirvana AI Governance: How AI Policymaking Is Committing Three Old Fallacies
von: Zhang, Jiawei
Veröffentlicht: (2024)
von: Zhang, Jiawei
Veröffentlicht: (2024)
STAR: SocioTechnical Approach to Red Teaming Language Models
von: Weidinger, Laura, et al.
Veröffentlicht: (2024)
von: Weidinger, Laura, et al.
Veröffentlicht: (2024)
Red Teaming LLMs as Socio-Technical Practice: From Exploration and Data Creation to Evaluation
von: Garcia, Adriana Alvarado, et al.
Veröffentlicht: (2026)
von: Garcia, Adriana Alvarado, et al.
Veröffentlicht: (2026)
Characterizing Scam-Driven Human Trafficking Across Chinese Borders and Online Community Responses on RedNote
von: Zheng, Jiamin, et al.
Veröffentlicht: (2026)
von: Zheng, Jiamin, et al.
Veröffentlicht: (2026)
The Situate AI Guidebook: Co-Designing a Toolkit to Support Multi-Stakeholder Early-stage Deliberations Around Public Sector AI Proposals
von: Kawakami, Anna, et al.
Veröffentlicht: (2024)
von: Kawakami, Anna, et al.
Veröffentlicht: (2024)
'Simulacrum of Stories': Examining Large Language Models as Qualitative Research Participants
von: Kapania, Shivani, et al.
Veröffentlicht: (2024)
von: Kapania, Shivani, et al.
Veröffentlicht: (2024)
Managing Project Teams in an Online Class of 1000+ Students
von: Anaraki, Nazanin Tabatabaei, et al.
Veröffentlicht: (2024)
von: Anaraki, Nazanin Tabatabaei, et al.
Veröffentlicht: (2024)
Generative AI in Medicine
von: Shanmugam, Divya, et al.
Veröffentlicht: (2024)
von: Shanmugam, Divya, et al.
Veröffentlicht: (2024)
Towards Standardizing AI Bias Exploration
von: Krasanakis, Emmanouil, et al.
Veröffentlicht: (2024)
von: Krasanakis, Emmanouil, et al.
Veröffentlicht: (2024)
Towards Human-AI Complementarity with Prediction Sets
von: De Toni, Giovanni, et al.
Veröffentlicht: (2024)
von: De Toni, Giovanni, et al.
Veröffentlicht: (2024)
Navigating Uncertainties: How GenAI Developers Document Their Models on Open-Source Platforms
von: Tang, Ningjing, et al.
Veröffentlicht: (2025)
von: Tang, Ningjing, et al.
Veröffentlicht: (2025)
Privacy Perspectives and Practices of Chinese Smart Home Product Teams
von: He, Shijing, et al.
Veröffentlicht: (2025)
von: He, Shijing, et al.
Veröffentlicht: (2025)
Human-Aligned Calibration for AI-Assisted Decision Making
von: Benz, Nina L. Corvelo, et al.
Veröffentlicht: (2023)
von: Benz, Nina L. Corvelo, et al.
Veröffentlicht: (2023)
Rethinking AI Evaluation in Education: The TEACH-AI Framework and Benchmark for Generative AI Assistants
von: Ding, Shi, et al.
Veröffentlicht: (2025)
von: Ding, Shi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
PersonaTeaming: Supporting Persona-Driven Red-Teaming for Generative AI
von: Deng, Wesley Hanwen, et al.
Veröffentlicht: (2026) -
The Generative AI Ethics Playbook
von: Smith, Jessie J., et al.
Veröffentlicht: (2024) -
On the Pros and Cons of Active Learning for Moral Preference Elicitation
von: Keswani, Vijay, et al.
Veröffentlicht: (2024) -
Not a Silver Bullet for Loneliness: How Attachment and Age Shape Intimacy with AI Companions
von: Ciriello, Raffaele, et al.
Veröffentlicht: (2026) -
Effective Automation to Support the Human Infrastructure in AI Red Teaming
von: Zhang, Alice Qian, et al.
Veröffentlicht: (2025)