Image Safeguarding: Reasoning with Conditional Vision Language Model and Obfuscating Unsafe Content Counterfactually
Fuente:
arXiv
Saved in:
| Main Authors: | Bethany, Mazal, Wherry, Brandon, Vishwamitra, Nishant, Najafirad, Peyman |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Deciphering Textual Authenticity: A Generalized Strategy through the Lens of Large Language Semantics for Detecting Human vs. Machine-Generated Text
by: Bethany, Mazal, et al.
Published: (2024)
by: Bethany, Mazal, et al.
Published: (2024)
Enhancing Event Reasoning in Large Language Models through Instruction Fine-Tuning with Semantic Causal Graphs
by: Bethany, Mazal, et al.
Published: (2024)
by: Bethany, Mazal, et al.
Published: (2024)
CAMOUFLAGE: Exploiting Misinformation Detection Systems Through LLM-driven Adversarial Claim Transformation
by: Bethany, Mazal, et al.
Published: (2025)
by: Bethany, Mazal, et al.
Published: (2025)
Lateral Phishing With Large Language Models: A Large Organization Comparative Study
by: Bethany, Mazal, et al.
Published: (2024)
by: Bethany, Mazal, et al.
Published: (2024)
Agentic Adversarial Rewriting Exposes Architectural Vulnerabilities in Black-Box NLP Pipelines
by: Bethany, Mazal, et al.
Published: (2026)
by: Bethany, Mazal, et al.
Published: (2026)
Can Reinforcement Learning Unlock the Hidden Dangers in Aligned Large Language Models?
by: Karkevandi, Mohammad Bahrami, et al.
Published: (2024)
by: Karkevandi, Mohammad Bahrami, et al.
Published: (2024)
FLAVARS: A Multimodal Foundational Language and Vision Alignment Model for Remote Sensing
by: Corley, Isaac, et al.
Published: (2025)
by: Corley, Isaac, et al.
Published: (2025)
Safe Vision-Language Models via Unsafe Weights Manipulation
by: D'Incà, Moreno, et al.
Published: (2025)
by: D'Incà, Moreno, et al.
Published: (2025)
Cross-Modal Obfuscation for Jailbreak Attacks on Large Vision-Language Models
by: Jiang, Lei, et al.
Published: (2025)
by: Jiang, Lei, et al.
Published: (2025)
Cultural Counterfactuals: Evaluating Cultural Biases in Large Vision-Language Models with Counterfactual Examples
by: Howard, Phillip, et al.
Published: (2026)
by: Howard, Phillip, et al.
Published: (2026)
LLaVAShield: Safeguarding Multimodal Multi-Turn Dialogues in Vision-Language Models
by: Huang, Guolei, et al.
Published: (2025)
by: Huang, Guolei, et al.
Published: (2025)
Estimating Earthquake Magnitude in Sentinel-1 Imagery via Ranking
by: Cambrin, Daniele Rege, et al.
Published: (2024)
by: Cambrin, Daniele Rege, et al.
Published: (2024)
OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
by: Wang, Shihao, et al.
Published: (2024)
by: Wang, Shihao, et al.
Published: (2024)
OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
by: Wang, Shihao, et al.
Published: (2025)
by: Wang, Shihao, et al.
Published: (2025)
Safeguarding Vision-Language Models: Mitigating Vulnerabilities to Gaussian Noise in Perturbation-based Attacks
by: Wang, Jiawei, et al.
Published: (2025)
by: Wang, Jiawei, et al.
Published: (2025)
Uncovering Bias in Large Vision-Language Models at Scale with Counterfactuals
by: Howard, Phillip, et al.
Published: (2024)
by: Howard, Phillip, et al.
Published: (2024)
Vision Language Modeling of Content, Distortion and Appearance for Image Quality Assessment
by: Zhou, Fei, et al.
Published: (2024)
by: Zhou, Fei, et al.
Published: (2024)
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
by: Yuan, Lingzhi, et al.
Published: (2025)
by: Yuan, Lingzhi, et al.
Published: (2025)
SSCR: Iterative Language-Based Image Editing via Self-Supervised Counterfactual Reasoning
by: Fu, Tsu-Jui, et al.
Published: (2020)
by: Fu, Tsu-Jui, et al.
Published: (2020)
Chartographer: Counterfactual Chart Generation for Evaluating Vision-Language Models
by: Jiang, Yifan, et al.
Published: (2026)
by: Jiang, Yifan, et al.
Published: (2026)
To Generate or Not? Safety-Driven Unlearned Diffusion Models Are Still Easy To Generate Unsafe Images ... For Now
by: Zhang, Yimeng, et al.
Published: (2023)
by: Zhang, Yimeng, et al.
Published: (2023)
Accelerating Conditional Prompt Learning via Masked Image Modeling for Vision-Language Models
by: Bui, Phuoc-Nguyen, et al.
Published: (2025)
by: Bui, Phuoc-Nguyen, et al.
Published: (2025)
Uncovering Bias in Large Vision-Language Models with Counterfactuals
by: Howard, Phillip, et al.
Published: (2024)
by: Howard, Phillip, et al.
Published: (2024)
Llama Guard 3 Vision: Safeguarding Human-AI Image Understanding Conversations
by: Chi, Jianfeng, et al.
Published: (2024)
by: Chi, Jianfeng, et al.
Published: (2024)
Counterfactual Vision-and-Language Navigation via Adversarial Path Sampling
by: Fu, Tsu-Jui, et al.
Published: (2019)
by: Fu, Tsu-Jui, et al.
Published: (2019)
Building Reasonable Inference for Vision-Language Models in Blind Image Quality Assessment
by: Li, Yuan, et al.
Published: (2025)
by: Li, Yuan, et al.
Published: (2025)
SC-Pro: Training-Free Framework for Defending Unsafe Image Synthesis Attack
by: Park, Junha, et al.
Published: (2025)
by: Park, Junha, et al.
Published: (2025)
SocialCounterfactuals: Probing and Mitigating Intersectional Social Biases in Vision-Language Models with Counterfactual Examples
by: Howard, Phillip, et al.
Published: (2023)
by: Howard, Phillip, et al.
Published: (2023)
Deep Generative Models Unveil Patterns in Medical Images Through Vision-Language Conditioning
by: Xing, Xiaodan, et al.
Published: (2024)
by: Xing, Xiaodan, et al.
Published: (2024)
Training-free Conditional Image Embedding Framework Leveraging Large Vision Language Models
by: Kawarada, Masayuki, et al.
Published: (2025)
by: Kawarada, Masayuki, et al.
Published: (2025)
UAV-VLN: End-to-End Vision Language guided Navigation for UAVs
by: Saxena, Pranav, et al.
Published: (2025)
by: Saxena, Pranav, et al.
Published: (2025)
Diagnosing and Repairing Unsafe Channels in Vision-Language Models via Causal Discovery and Dual-Modal Safety Subspace Projection
by: Fu, Jinhu, et al.
Published: (2026)
by: Fu, Jinhu, et al.
Published: (2026)
Jailbreaking Large Language Models with Symbolic Mathematics
by: Bethany, Emet, et al.
Published: (2024)
by: Bethany, Emet, et al.
Published: (2024)
Recognition through Reasoning: Reinforcing Image Geo-localization with Large Vision-Language Models
by: Li, Ling, et al.
Published: (2025)
by: Li, Ling, et al.
Published: (2025)
Attention Shift: Steering AI Away from Unsafe Content
by: Garg, Shivank, et al.
Published: (2024)
by: Garg, Shivank, et al.
Published: (2024)
Seeing the roads through the trees: A benchmark for modeling spatial dependencies with aerial imagery
by: Robinson, Caleb, et al.
Published: (2024)
by: Robinson, Caleb, et al.
Published: (2024)
Streamlining Image Editing with Layered Diffusion Brushes
by: Gholami, Peyman, et al.
Published: (2024)
by: Gholami, Peyman, et al.
Published: (2024)
Counterfactual Stress Testing for Image Classification Models
by: Stammel, Moritz, et al.
Published: (2026)
by: Stammel, Moritz, et al.
Published: (2026)
GeoShield: Safeguarding Geolocation Privacy from Vision-Language Models via Adversarial Perturbations
by: Liu, Xinwei, et al.
Published: (2025)
by: Liu, Xinwei, et al.
Published: (2025)
Explanation-Driven Counterfactual Testing for Faithfulness in Vision-Language Model Explanations
by: Ding, Sihao, et al.
Published: (2025)
by: Ding, Sihao, et al.
Published: (2025)
Similar Items
-
Deciphering Textual Authenticity: A Generalized Strategy through the Lens of Large Language Semantics for Detecting Human vs. Machine-Generated Text
by: Bethany, Mazal, et al.
Published: (2024) -
Enhancing Event Reasoning in Large Language Models through Instruction Fine-Tuning with Semantic Causal Graphs
by: Bethany, Mazal, et al.
Published: (2024) -
CAMOUFLAGE: Exploiting Misinformation Detection Systems Through LLM-driven Adversarial Claim Transformation
by: Bethany, Mazal, et al.
Published: (2025) -
Lateral Phishing With Large Language Models: A Large Organization Comparative Study
by: Bethany, Mazal, et al.
Published: (2024) -
Agentic Adversarial Rewriting Exposes Architectural Vulnerabilities in Black-Box NLP Pipelines
by: Bethany, Mazal, et al.
Published: (2026)