SafetyPairs: Isolating Safety Critical Image Features with Counterfactual Image Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Helbling, Alec, Palaskar, Shruti, Krishna, Kundan, Chau, Polo, Gatys, Leon, Cheng, Joseph Yitan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ClickDiffusion: Harnessing LLMs for Interactive Precise Image Editing
von: Helbling, Alec, et al.
Veröffentlicht: (2024)
von: Helbling, Alec, et al.
Veröffentlicht: (2024)
VLSU: Mapping the Limits of Joint Multimodal Understanding for AI Safety
von: Palaskar, Shruti, et al.
Veröffentlicht: (2025)
von: Palaskar, Shruti, et al.
Veröffentlicht: (2025)
ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features
von: Helbling, Alec, et al.
Veröffentlicht: (2025)
von: Helbling, Alec, et al.
Veröffentlicht: (2025)
CounterScene: Counterfactual Causal Reasoning in Generative World Models for Safety-Critical Closed-Loop Evaluation
von: Jing, Bowen, et al.
Veröffentlicht: (2026)
von: Jing, Bowen, et al.
Veröffentlicht: (2026)
Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment
von: Krishna, Kundan, et al.
Veröffentlicht: (2025)
von: Krishna, Kundan, et al.
Veröffentlicht: (2025)
Benchmarking Counterfactual Image Generation
von: Melistas, Thomas, et al.
Veröffentlicht: (2024)
von: Melistas, Thomas, et al.
Veröffentlicht: (2024)
iSafetyBench: A video-language benchmark for safety in industrial environment
von: Abdullah, Raiyaan, et al.
Veröffentlicht: (2025)
von: Abdullah, Raiyaan, et al.
Veröffentlicht: (2025)
Perceptual Classifiers: Detecting Generative Images using Perceptual Features
von: Durbha, Krishna Srikar, et al.
Veröffentlicht: (2025)
von: Durbha, Krishna Srikar, et al.
Veröffentlicht: (2025)
SafeGen-Bench: Benchmarking Safety in Image-Conditioned Text-to-Video Generation
von: Ma, Yingzi, et al.
Veröffentlicht: (2026)
von: Ma, Yingzi, et al.
Veröffentlicht: (2026)
The Path to Reconciling Quality and Safety in Text-to-Image Generation: Dataset, Method, and Evaluation
von: Ruan, Shouwei, et al.
Veröffentlicht: (2025)
von: Ruan, Shouwei, et al.
Veröffentlicht: (2025)
Image Classification for Snow Detection to Improve Pedestrian Safety
von: de Deijn, Ricardo, et al.
Veröffentlicht: (2024)
von: de Deijn, Ricardo, et al.
Veröffentlicht: (2024)
Safety of Multimodal Large Language Models on Images and Texts
von: Liu, Xin, et al.
Veröffentlicht: (2024)
von: Liu, Xin, et al.
Veröffentlicht: (2024)
To Generate or Not? Safety-Driven Unlearned Diffusion Models Are Still Easy To Generate Unsafe Images ... For Now
von: Zhang, Yimeng, et al.
Veröffentlicht: (2023)
von: Zhang, Yimeng, et al.
Veröffentlicht: (2023)
Adversarial Generation and Collaborative Evolution of Safety-Critical Scenarios for Autonomous Vehicles
von: Liu, Jiangfan, et al.
Veröffentlicht: (2025)
von: Liu, Jiangfan, et al.
Veröffentlicht: (2025)
Enhancing Counterfactual Image Generation Using Mahalanobis Distance with Distribution Preferences in Feature Space
von: Zhang, Yukai, et al.
Veröffentlicht: (2024)
von: Zhang, Yukai, et al.
Veröffentlicht: (2024)
Point and Instruct: Enabling Precise Image Editing by Unifying Direct Manipulation and Text Instructions
von: Helbling, Alec, et al.
Veröffentlicht: (2024)
von: Helbling, Alec, et al.
Veröffentlicht: (2024)
RenderBender: A Survey on Adversarial Attacks Using Differentiable Rendering
von: Hull, Matthew, et al.
Veröffentlicht: (2024)
von: Hull, Matthew, et al.
Veröffentlicht: (2024)
Cycle Diffusion Model for Counterfactual Image Generation
von: Huang, Fangrui, et al.
Veröffentlicht: (2025)
von: Huang, Fangrui, et al.
Veröffentlicht: (2025)
When Safety Collides: Resolving Multi-Category Harmful Conflicts in Text-to-Image Diffusion via Adaptive Safety Guidance
von: Xiang, Yongli, et al.
Veröffentlicht: (2026)
von: Xiang, Yongli, et al.
Veröffentlicht: (2026)
Unifying Image Counterfactuals and Feature Attributions with Latent-Space Adversarial Attacks
von: Goldwasser, Jeremy, et al.
Veröffentlicht: (2025)
von: Goldwasser, Jeremy, et al.
Veröffentlicht: (2025)
Beyond the Safety Tax: Mitigating Unsafe Text-to-Image Generation via External Safety Rectification
von: Meng, Xiangtao, et al.
Veröffentlicht: (2025)
von: Meng, Xiangtao, et al.
Veröffentlicht: (2025)
Counterfactual Image Editing
von: Pan, Yushu, et al.
Veröffentlicht: (2024)
von: Pan, Yushu, et al.
Veröffentlicht: (2024)
Learning an Image Editing Model without Image Editing Pairs
von: Kumari, Nupur, et al.
Veröffentlicht: (2025)
von: Kumari, Nupur, et al.
Veröffentlicht: (2025)
Low-Effort Jailbreak Attacks Against Text-to-Image Safety Filters
von: Mustafa, Ahmed B, et al.
Veröffentlicht: (2026)
von: Mustafa, Ahmed B, et al.
Veröffentlicht: (2026)
MiSCHiEF: A Benchmark in Minimal-Pairs of Safety and Culture for Holistic Evaluation of Fine-Grained Image-Caption Alignment
von: Banerjee, Sagarika, et al.
Veröffentlicht: (2026)
von: Banerjee, Sagarika, et al.
Veröffentlicht: (2026)
Counterfactual Stress Testing for Image Classification Models
von: Stammel, Moritz, et al.
Veröffentlicht: (2026)
von: Stammel, Moritz, et al.
Veröffentlicht: (2026)
SafeMVDrive: Multi-view Safety-Critical Driving Video Synthesis in the Real World Domain
von: Zhou, Jiawei, et al.
Veröffentlicht: (2025)
von: Zhou, Jiawei, et al.
Veröffentlicht: (2025)
AlignGuard: Scalable Safety Alignment for Text-to-Image Generation
von: Liu, Runtao, et al.
Veröffentlicht: (2024)
von: Liu, Runtao, et al.
Veröffentlicht: (2024)
Beyond Strict Pairing: Arbitrarily Paired Training for High-Performance Infrared and Visible Image Fusion
von: Deng, Yanglin, et al.
Veröffentlicht: (2026)
von: Deng, Yanglin, et al.
Veröffentlicht: (2026)
ELITE: Enhanced Language-Image Toxicity Evaluation for Safety
von: Lee, Wonjun, et al.
Veröffentlicht: (2025)
von: Lee, Wonjun, et al.
Veröffentlicht: (2025)
Personalized Safety Alignment for Text-to-Image Diffusion Models
von: Lei, Yu, et al.
Veröffentlicht: (2025)
von: Lei, Yu, et al.
Veröffentlicht: (2025)
Bridging Foundation Models and ASTM Metallurgical Standards for Automated Grain Size Estimation from Microscopy Images
von: Mueez, Abdul, et al.
Veröffentlicht: (2026)
von: Mueez, Abdul, et al.
Veröffentlicht: (2026)
Eyes Closed, Safety On: Protecting Multimodal LLMs via Image-to-Text Transformation
von: Gou, Yunhao, et al.
Veröffentlicht: (2024)
von: Gou, Yunhao, et al.
Veröffentlicht: (2024)
SurrogatePrompt: Bypassing the Safety Filter of Text-to-Image Models via Substitution
von: Ba, Zhongjie, et al.
Veröffentlicht: (2023)
von: Ba, Zhongjie, et al.
Veröffentlicht: (2023)
Diffusion Counterfactuals for Image Regressors
von: Ha, Trung Duc, et al.
Veröffentlicht: (2025)
von: Ha, Trung Duc, et al.
Veröffentlicht: (2025)
USIGAN: Unbalanced Self-Information Feature Transport for Weakly Paired Image IHC Virtual Staining
von: Peng, Yue, et al.
Veröffentlicht: (2025)
von: Peng, Yue, et al.
Veröffentlicht: (2025)
ThermalDiffusion: Visual-to-Thermal Image-to-Image Translation for Autonomous Navigation
von: Bansal, Shruti, et al.
Veröffentlicht: (2025)
von: Bansal, Shruti, et al.
Veröffentlicht: (2025)
An Interpretable Local Editing Model for Counterfactual Medical Image Generation
von: Min, Hyungi, et al.
Veröffentlicht: (2026)
von: Min, Hyungi, et al.
Veröffentlicht: (2026)
Boosting Generative Image Modeling via Joint Image-Feature Synthesis
von: Kouzelis, Theodoros, et al.
Veröffentlicht: (2025)
von: Kouzelis, Theodoros, et al.
Veröffentlicht: (2025)
Visual Polarization Measurement Using Counterfactual Image Generation
von: Mosaffa, Mohammad, et al.
Veröffentlicht: (2025)
von: Mosaffa, Mohammad, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ClickDiffusion: Harnessing LLMs for Interactive Precise Image Editing
von: Helbling, Alec, et al.
Veröffentlicht: (2024) -
VLSU: Mapping the Limits of Joint Multimodal Understanding for AI Safety
von: Palaskar, Shruti, et al.
Veröffentlicht: (2025) -
ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features
von: Helbling, Alec, et al.
Veröffentlicht: (2025) -
CounterScene: Counterfactual Causal Reasoning in Generative World Models for Safety-Critical Closed-Loop Evaluation
von: Jing, Bowen, et al.
Veröffentlicht: (2026) -
Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment
von: Krishna, Kundan, et al.
Veröffentlicht: (2025)