Adversarial Nibbler: An Open Red-Teaming Method for Identifying Diverse Harms in Text-to-Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Quaye, Jessica, Parrish, Alicia, Inel, Oana, Rastogi, Charvi, Kirk, Hannah Rose, Kahng, Minsuk, van Liemt, Erin, Bartolo, Max, Tsang, Jess, White, Justin, Clement, Nathan, Mosquera, Rafael, Ciro, Juan, Reddi, Vijay Janapa, Aroyo, Lora |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image Models
by: Quaye, Jessica, et al.
Published: (2025)
by: Quaye, Jessica, et al.
Published: (2025)
Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South
by: Rastogi, Charvi, et al.
Published: (2026)
by: Rastogi, Charvi, et al.
Published: (2026)
TinyTorch: Building Machine Learning Systems from First Principles
by: Reddi, Vijay Janapa
Published: (2026)
by: Reddi, Vijay Janapa
Published: (2026)
Generative AI Agents in Autonomous Machines: A Safety Perspective
by: Jabbour, Jason, et al.
Published: (2024)
by: Jabbour, Jason, et al.
Published: (2024)
SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early Alignment
by: Jeung, Wonje, et al.
Published: (2025)
by: Jeung, Wonje, et al.
Published: (2025)
The Magnificent Seven Challenges and Opportunities in Domain-Specific Accelerator Design for Autonomous Systems
by: Neuman, Sabrina M., et al.
Published: (2024)
by: Neuman, Sabrina M., et al.
Published: (2024)
Data-Prompt Co-Evolution: Growing Test Sets to Refine LLM Behavior
by: Lee, Minjae, et al.
Published: (2025)
by: Lee, Minjae, et al.
Published: (2025)
FedStaleWeight: Buffered Asynchronous Federated Learning with Fair Aggregation via Staleness Reweighting
by: Ma, Jeffrey, et al.
Published: (2024)
by: Ma, Jeffrey, et al.
Published: (2024)
Decoding Safety Feedback from Diverse Raters: A Data-driven Lens on Responsiveness to Severity
by: Mishra, Pushkar, et al.
Published: (2025)
by: Mishra, Pushkar, et al.
Published: (2025)
"Just a strange pic": Evaluating 'safety' in GenAI Image safety annotation tasks from diverse annotators' perspectives
by: Wang, Ding, et al.
Published: (2025)
by: Wang, Ding, et al.
Published: (2025)
Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models
by: Rastogi, Charvi, et al.
Published: (2025)
by: Rastogi, Charvi, et al.
Published: (2025)
Insights on Disagreement Patterns in Multimodal Safety Perception across Diverse Rater Groups
by: Rastogi, Charvi, et al.
Published: (2024)
by: Rastogi, Charvi, et al.
Published: (2024)
Understanding the Dataset Practitioners Behind Large Language Model Development
by: Qian, Crystal, et al.
Published: (2024)
by: Qian, Crystal, et al.
Published: (2024)
VLSlice: Interactive Vision-and-Language Slice Discovery
by: Slyman, Eric, et al.
Published: (2023)
by: Slyman, Eric, et al.
Published: (2023)
SocratiQ: A Generative AI-Powered Learning Companion for Personalized Education and Broader Accessibility
by: Jabbour, Jason, et al.
Published: (2025)
by: Jabbour, Jason, et al.
Published: (2025)
Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language Models
by: Ibrahim, Lujain, et al.
Published: (2025)
by: Ibrahim, Lujain, et al.
Published: (2025)
Tabula: Efficiently Computing Nonlinear Activation Functions for Secure Neural Network Inference
by: Lam, Maximilian, et al.
Published: (2022)
by: Lam, Maximilian, et al.
Published: (2022)
Automatic Histograms: Leveraging Language Models for Text Dataset Exploration
by: Reif, Emily, et al.
Published: (2024)
by: Reif, Emily, et al.
Published: (2024)
When Cars Have Stereotypes: Auditing Demographic Bias in Objects from Text-to-Image Models
by: Choi, Dasol, et al.
Published: (2025)
by: Choi, Dasol, et al.
Published: (2025)
Whom do Explanations Serve? A Systematic Literature Survey of User Characteristics in Explainable Recommender Systems Evaluation
by: Wardatzky, Kathrin, et al.
Published: (2024)
by: Wardatzky, Kathrin, et al.
Published: (2024)
D-RDW: Diversity-Driven Random Walks for News Recommender Systems
by: Li, Runze, et al.
Published: (2025)
by: Li, Runze, et al.
Published: (2025)
Informfully Recommenders -- Reproducibility Framework for Diversity-aware Intra-session Recommendations
by: Heitz, Lucien, et al.
Published: (2025)
by: Heitz, Lucien, et al.
Published: (2025)
TinyML Security: Exploring Vulnerabilities in Resource-Constrained Machine Learning Systems
by: Huckelberry, Jacob, et al.
Published: (2024)
by: Huckelberry, Jacob, et al.
Published: (2024)
Multi-Agent Reinforcement Learning for Sample-Efficient Deep Neural Network Mapping
by: Krishnan, Srivatsan, et al.
Published: (2025)
by: Krishnan, Srivatsan, et al.
Published: (2025)
Adaptive Surrogate Gradients for Sequential Reinforcement Learning in Spiking Neural Networks
by: Berghe, Korneel Van den, et al.
Published: (2025)
by: Berghe, Korneel Van den, et al.
Published: (2025)
Materiality and Risk in the Age of Pervasive AI Sensors
by: Sloane, Mona, et al.
Published: (2024)
by: Sloane, Mona, et al.
Published: (2024)
Evaluating Language Models for Harmful Manipulation
by: Akbulut, Canfer, et al.
Published: (2026)
by: Akbulut, Canfer, et al.
Published: (2026)
From Perception to Decision: Assessing the Role of Chart Types Affordances in High-Level Decision Tasks
by: Li, Yixuan, et al.
Published: (2024)
by: Li, Yixuan, et al.
Published: (2024)
Who Defines "Best"? Towards Interactive, User-Defined Evaluation of LLM Leaderboards
by: Jung, Minji, et al.
Published: (2026)
by: Jung, Minji, et al.
Published: (2026)
Slm-mux: Orchestrating small language models for reasoning
by: Wang, Chenyu, et al.
Published: (2025)
by: Wang, Chenyu, et al.
Published: (2025)
SwizzlePerf: Hardware-Aware LLMs for GPU Kernel Performance Optimization
by: Tschand, Arya, et al.
Published: (2025)
by: Tschand, Arya, et al.
Published: (2025)
Aligning Object Detector Bounding Boxes with Human Preference
by: Strafforello, Ombretta, et al.
Published: (2024)
by: Strafforello, Ombretta, et al.
Published: (2024)
GRASP: A Disagreement Analysis Framework to Assess Group Associations in Perspectives
by: Prabhakaran, Vinodkumar, et al.
Published: (2023)
by: Prabhakaran, Vinodkumar, et al.
Published: (2023)
Paradoxical Vitiligo Induced by Adalimumab in Ankylosing Spondylitis With Uveitis: Clinical Management and Therapeutic Considerations
by: Tuba Yuce Inel
Published: (2026)
by: Tuba Yuce Inel
Published: (2026)
SWE-fficiency: Can Language Models Optimize Real-World Repositories on Real Workloads?
by: Ma, Jeffrey Jian, et al.
Published: (2025)
by: Ma, Jeffrey Jian, et al.
Published: (2025)
Ontology of Belief Diversity: A Community-Based Epistemological Approach
by: Fischella, Tyler, et al.
Published: (2024)
by: Fischella, Tyler, et al.
Published: (2024)
Interactive Prompt Debugging with Sequence Salience
by: Tenney, Ian, et al.
Published: (2024)
by: Tenney, Ian, et al.
Published: (2024)
DWARF: Disease-weighted network for attention map refinement
by: Luo, Haozhe, et al.
Published: (2024)
by: Luo, Haozhe, et al.
Published: (2024)
Exploring first‐year occupational therapy students' perspectives of an On‐Country experience: A study from an Australian undergraduate program
by: Kieva Richards, et al.
Published: (2025)
by: Kieva Richards, et al.
Published: (2025)
COSMIC: Enabling Full-Stack Co-Design and Optimization of Distributed Machine Learning Systems
by: Raju, Aditi, et al.
Published: (2025)
by: Raju, Aditi, et al.
Published: (2025)
Similar Items
-
From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image Models
by: Quaye, Jessica, et al.
Published: (2025) -
Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South
by: Rastogi, Charvi, et al.
Published: (2026) -
TinyTorch: Building Machine Learning Systems from First Principles
by: Reddi, Vijay Janapa
Published: (2026) -
Generative AI Agents in Autonomous Machines: A Safety Perspective
by: Jabbour, Jason, et al.
Published: (2024) -
SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early Alignment
by: Jeung, Wonje, et al.
Published: (2025)