When Image Generation Goes Wrong: A Safety Analysis of Stable Diffusion Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Schneider, Matthias, Hagendorff, Thilo |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Mapping the Ethics of Generative AI: A Comprehensive Scoping Review
par: Hagendorff, Thilo
Publié: (2024)
par: Hagendorff, Thilo
Publié: (2024)
On the Inevitability of Left-Leaning Political Bias in Aligned Language Models
par: Hagendorff, Thilo
Publié: (2025)
par: Hagendorff, Thilo
Publié: (2025)
PRIDE -- Parameter-Efficient Reduction of Identity Discrimination for Equality in LLMs
par: Menke, Maluna, et autres
Publié: (2025)
par: Menke, Maluna, et autres
Publié: (2025)
Evaluation Awareness in Language Models Has Limited Effect on Behaviour
par: Knecht, Amelie, et autres
Publié: (2026)
par: Knecht, Amelie, et autres
Publié: (2026)
Fairness Hacking: The Malicious Practice of Shrouding Unfairness in Algorithms
par: Meding, Kristof, et autres
Publié: (2023)
par: Meding, Kristof, et autres
Publié: (2023)
Compromising Honesty and Harmlessness in Language Models via Deception Attacks
par: Vaugrante, Laurène, et autres
Publié: (2025)
par: Vaugrante, Laurène, et autres
Publié: (2025)
Speciesism in AI: Evaluating Discrimination Against Animals in Large Language Models
par: Jotautaitė, Monika, et autres
Publié: (2025)
par: Jotautaitė, Monika, et autres
Publié: (2025)
Deception Abilities Emerged in Large Language Models
par: Hagendorff, Thilo
Publié: (2023)
par: Hagendorff, Thilo
Publié: (2023)
Beyond Chains of Thought: Benchmarking Latent-Space Reasoning Abilities in Large Language Models
par: Hagendorff, Thilo, et autres
Publié: (2025)
par: Hagendorff, Thilo, et autres
Publié: (2025)
When Handwriting Goes Social: Creativity, Anonymity, and Communication in Graphonymous Online Spaces
par: Purohit, Aditya Kumar, et autres
Publié: (2026)
par: Purohit, Aditya Kumar, et autres
Publié: (2026)
When Machines Get It Wrong: Large Language Models Perpetuate Autism Myths More Than Humans Do
par: Garrido-Merchán, Eduardo C., et autres
Publié: (2026)
par: Garrido-Merchán, Eduardo C., et autres
Publié: (2026)
When Abundance Goes Wrong
par: Lee Skallerup Bessette
Publié: (2024)
par: Lee Skallerup Bessette
Publié: (2024)
A Looming Replication Crisis in Evaluating Behavior in Language Models? Evidence and Solutions
par: Vaugrante, Laurène, et autres
Publié: (2024)
par: Vaugrante, Laurène, et autres
Publié: (2024)
Emergently Misaligned Language Models Show Behavioral Self-Awareness That Shifts With Subsequent Realignment
par: Vaugrante, Laurène, et autres
Publié: (2026)
par: Vaugrante, Laurène, et autres
Publié: (2026)
Large Reasoning Models Are Autonomous Jailbreak Agents
par: Hagendorff, Thilo, et autres
Publié: (2025)
par: Hagendorff, Thilo, et autres
Publié: (2025)
A.I. In All The Wrong Places
par: Böhlen, Marc, et autres
Publié: (2024)
par: Böhlen, Marc, et autres
Publié: (2024)
When Information Abundance Goes Wrong
par: Lee Skallerup Bessette
Publié: (2024)
par: Lee Skallerup Bessette
Publié: (2024)
When Openness Fails: Lessons from System Safety for Assessing Openness in AI
par: Paris, Tamara, et autres
Publié: (2025)
par: Paris, Tamara, et autres
Publié: (2025)
AI Safety in Generative AI Large Language Models: A Survey
par: Chua, Jaymari, et autres
Publié: (2024)
par: Chua, Jaymari, et autres
Publié: (2024)
All Models Are Wrong, But Can They Be Useful? Lessons from COVID-19 Agent-Based Models: A Systematic Review
par: Von Hoene, Emma, et autres
Publié: (2025)
par: Von Hoene, Emma, et autres
Publié: (2025)
Probabilistic Analysis of Copyright Disputes and Generative AI Safety
par: Chiba-Okabe, Hiroaki
Publié: (2024)
par: Chiba-Okabe, Hiroaki
Publié: (2024)
Unmasking the Canvas: A Dynamic Benchmark for Image Generation Jailbreaking and LLM Content Safety
par: Nair, Variath Madhupal Gautham, et autres
Publié: (2025)
par: Nair, Variath Madhupal Gautham, et autres
Publié: (2025)
Descriptions of women are longer than that of men: An analysis of gender portrayal prompts in Stable Diffusion
par: Asadchy, Yan, et autres
Publié: (2024)
par: Asadchy, Yan, et autres
Publié: (2024)
A Large Scale Analysis of Gender Biases in Text-to-Image Generative Models
par: Girrbach, Leander, et autres
Publié: (2025)
par: Girrbach, Leander, et autres
Publié: (2025)
Do Small Language Models Know When They're Wrong? Confidence-Based Cascade Scoring for Educational Assessment
par: Burleigh, Tyler
Publié: (2026)
par: Burleigh, Tyler
Publié: (2026)
Persuasion and Safety in the Era of Generative AI
par: Kong, Haein
Publié: (2025)
par: Kong, Haein
Publié: (2025)
Diffusion-driven pattern formation in an opinion dynamical network model
par: Mauch, Tim, et autres
Publié: (2025)
par: Mauch, Tim, et autres
Publié: (2025)
Energy Scaling Laws for Diffusion Models: Quantifying Compute in Image Generation
par: Iyengar, Aniketh, et autres
Publié: (2025)
par: Iyengar, Aniketh, et autres
Publié: (2025)
A Task-Driven Human-AI Collaboration: When to Automate, When to Collaborate, When to Challenge
par: Afroogh, Saleh, et autres
Publié: (2025)
par: Afroogh, Saleh, et autres
Publié: (2025)
Debiasing Diffusion Model: Enhancing Fairness through Latent Representation Learning in Stable Diffusion Model
par: Huang, Lin-Chun, et autres
Publié: (2025)
par: Huang, Lin-Chun, et autres
Publié: (2025)
When Models Disagree: Rethinking LLM Evaluation for Public Comment Analysis
par: Najera, Aisha, et autres
Publié: (2026)
par: Najera, Aisha, et autres
Publié: (2026)
"Dark Triad" Model Organisms of Misalignment: Narrow Fine-Tuning Mirrors Human Antisocial Behavior
par: Lulla, Roshni, et autres
Publié: (2026)
par: Lulla, Roshni, et autres
Publié: (2026)
How to Assess Trustworthy AI in Practice
par: Zicari, Roberto V., et autres
Publié: (2022)
par: Zicari, Roberto V., et autres
Publié: (2022)
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
par: Zhang, Chi, et autres
Publié: (2025)
par: Zhang, Chi, et autres
Publié: (2025)
AI Meets the Classroom: When Do Large Language Models Harm Learning?
par: Lehmann, Matthias, et autres
Publié: (2024)
par: Lehmann, Matthias, et autres
Publié: (2024)
When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
par: Xiao, Yuxin, et autres
Publié: (2025)
par: Xiao, Yuxin, et autres
Publié: (2025)
Stable Signer: Hierarchical Sign Language Generative Model
par: Fang, Sen, et autres
Publié: (2025)
par: Fang, Sen, et autres
Publié: (2025)
Agentic Microphysics: A Manifesto for Generative AI Safety
par: Pierucci, Federico, et autres
Publié: (2026)
par: Pierucci, Federico, et autres
Publié: (2026)
When Attention Becomes Exposure in Generative Search
par: Alipour, Shayan, et autres
Publié: (2026)
par: Alipour, Shayan, et autres
Publié: (2026)
Seeking Human Security Consensus: A Unified Value Scale for Generative AI Value Safety
par: He, Ying, et autres
Publié: (2026)
par: He, Ying, et autres
Publié: (2026)
Documents similaires
-
Mapping the Ethics of Generative AI: A Comprehensive Scoping Review
par: Hagendorff, Thilo
Publié: (2024) -
On the Inevitability of Left-Leaning Political Bias in Aligned Language Models
par: Hagendorff, Thilo
Publié: (2025) -
PRIDE -- Parameter-Efficient Reduction of Identity Discrimination for Equality in LLMs
par: Menke, Maluna, et autres
Publié: (2025) -
Evaluation Awareness in Language Models Has Limited Effect on Behaviour
par: Knecht, Amelie, et autres
Publié: (2026) -
Fairness Hacking: The Malicious Practice of Shrouding Unfairness in Algorithms
par: Meding, Kristof, et autres
Publié: (2023)