SAGE: Exploring the Boundaries of Unsafe Concept Domain with Semantic-Augment Erasing
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Hongguang, Wei, Yunchao, Wang, Mengyu, Jiao, Siyu, Fang, Yan, Huang, Jiannan, Zhao, Yao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Do Concept Replacement Techniques Really Erase Unacceptable Concepts?
by: Das, Anudeep, et al.
Published: (2025)
by: Das, Anudeep, et al.
Published: (2025)
Bi-Erasing: A Bidirectional Framework for Concept Removal in Diffusion Models
by: Chen, Hao, et al.
Published: (2025)
by: Chen, Hao, et al.
Published: (2025)
Sparse Autoencoder as a Zero-Shot Classifier for Concept Erasing in Text-to-Image Diffusion Models
by: Tian, Zhihua, et al.
Published: (2025)
by: Tian, Zhihua, et al.
Published: (2025)
Towards Understanding Unsafe Video Generation
by: Pang, Yan, et al.
Published: (2024)
by: Pang, Yan, et al.
Published: (2024)
Collaborative Vision-Text Representation Optimizing for Open-Vocabulary Segmentation
by: Jiao, Siyu, et al.
Published: (2024)
by: Jiao, Siyu, et al.
Published: (2024)
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
by: Yuan, Lingzhi, et al.
Published: (2025)
by: Yuan, Lingzhi, et al.
Published: (2025)
Probing the Safety Response Boundary of Large Language Models via Unsafe Decoding Path Generation
by: Wang, Haoyu, et al.
Published: (2024)
by: Wang, Haoyu, et al.
Published: (2024)
Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities
by: Qu, Yiting, et al.
Published: (2025)
by: Qu, Yiting, et al.
Published: (2025)
Guarding the Gate: ConceptGuard Battles Concept-Level Backdoors in Concept Bottleneck Models
by: Lai, Songning, et al.
Published: (2024)
by: Lai, Songning, et al.
Published: (2024)
Attention Shift: Steering AI Away from Unsafe Content
by: Garg, Shivank, et al.
Published: (2024)
by: Garg, Shivank, et al.
Published: (2024)
Blind Spot Navigation: Evolutionary Discovery of Sensitive Semantic Concepts for LVLMs
by: Pan, Zihao, et al.
Published: (2025)
by: Pan, Zihao, et al.
Published: (2025)
Erased but Not Forgotten: How Backdoors Compromise Concept Erasure
by: Braun, Tobias, et al.
Published: (2025)
by: Braun, Tobias, et al.
Published: (2025)
SafeFFI: Efficient Sanitization at the Boundary Between Safe and Unsafe Code in Rust and Mixed-Language Applications
by: Braunsdorf, Oliver, et al.
Published: (2025)
by: Braunsdorf, Oliver, et al.
Published: (2025)
BlindU: Blind Machine Unlearning without Revealing Erasing Data
by: Wang, Weiqi, et al.
Published: (2026)
by: Wang, Weiqi, et al.
Published: (2026)
Targeted Fuzzing for Unsafe Rust Code: Leveraging Selective Instrumentation
by: Paaßen, David, et al.
Published: (2025)
by: Paaßen, David, et al.
Published: (2025)
CAT: Concept-level backdoor ATtacks for Concept Bottleneck Models
by: Lai, Songning, et al.
Published: (2024)
by: Lai, Songning, et al.
Published: (2024)
GuardTrace-VL: Detecting Unsafe Multimodel Reasoning via Iterative Safety Supervision
by: Xiang, Yuxiao, et al.
Published: (2025)
by: Xiang, Yuxiao, et al.
Published: (2025)
Friend or Foe Inside? Exploring In-Process Isolation to Maintain Memory Safety for Unsafe Rust
by: Gülmez, Merve, et al.
Published: (2023)
by: Gülmez, Merve, et al.
Published: (2023)
A Privacy-Preserving Semantic-Segmentation Method Using Domain-Adaptation Technique
by: Sueyoshi, Homare, et al.
Published: (2025)
by: Sueyoshi, Homare, et al.
Published: (2025)
Erasing Radio Frequency Fingerprints via Active Adversarial Perturbation
by: Lu, Zhaoyi, et al.
Published: (2024)
by: Lu, Zhaoyi, et al.
Published: (2024)
Beyond the Safety Tax: Mitigating Unsafe Text-to-Image Generation via External Safety Rectification
by: Meng, Xiangtao, et al.
Published: (2025)
by: Meng, Xiangtao, et al.
Published: (2025)
Random Erasing vs. Model Inversion: A Promising Defense or a False Hope?
by: Tran, Viet-Hung, et al.
Published: (2024)
by: Tran, Viet-Hung, et al.
Published: (2024)
Backdooring CLIP through Concept Confusion
by: Hu, Lijie, et al.
Published: (2025)
by: Hu, Lijie, et al.
Published: (2025)
Detecting Malicious Concepts without Image Generation in AI-Generated Content (AIGC)
by: Xu, Kun, et al.
Published: (2025)
by: Xu, Kun, et al.
Published: (2025)
Propagating Unsafe Actions in LLM Controlled Multi-Robot Collaboration via Single Robot Compromise
by: Huang, Zhen, et al.
Published: (2026)
by: Huang, Zhen, et al.
Published: (2026)
HomeSafe-Bench: Evaluating Vision-Language Models on Unsafe Action Detection for Embodied Agents in Household Scenarios
by: Pu, Jiayue, et al.
Published: (2026)
by: Pu, Jiayue, et al.
Published: (2026)
UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images
by: Qu, Yiting, et al.
Published: (2024)
by: Qu, Yiting, et al.
Published: (2024)
From Stealthy Data Fabrication to Unsafe Driving: Realistic Scenario Attacks on Collaborative Perception
by: Zhang, Qingzhao, et al.
Published: (2026)
by: Zhang, Qingzhao, et al.
Published: (2026)
Fast Summary-based Whole-program Analysis to Identify Unsafe Memory Accesses in Rust
by: Zhou, Jie, et al.
Published: (2023)
by: Zhou, Jie, et al.
Published: (2023)
On the Proactive Generation of Unsafe Images From Text-To-Image Models Using Benign Prompts
by: Wu, Yixin, et al.
Published: (2023)
by: Wu, Yixin, et al.
Published: (2023)
SAGE: Signal-Amplified Guided Embeddings for LLM-based Vulnerability Detection
by: Shan, Zhengyang, et al.
Published: (2026)
by: Shan, Zhengyang, et al.
Published: (2026)
Erasing Self-Supervised Learning Backdoor by Cluster Activation Masking
by: Qian, Shengsheng, et al.
Published: (2023)
by: Qian, Shengsheng, et al.
Published: (2023)
SandCell: Sandboxing Rust Beyond Unsafe Code
by: Zhang, Jialun, et al.
Published: (2025)
by: Zhang, Jialun, et al.
Published: (2025)
FASR: Automated Identification of Unsafe Control Actions in STPA
by: Dardik, Ian, et al.
Published: (2026)
by: Dardik, Ian, et al.
Published: (2026)
Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM Security
by: Fang, Xiang, et al.
Published: (2026)
by: Fang, Xiang, et al.
Published: (2026)
E-SAGE: Explainability-based Defense Against Backdoor Attacks on Graph Neural Networks
by: Yuan, Dingqiang, et al.
Published: (2024)
by: Yuan, Dingqiang, et al.
Published: (2024)
SAGE: Sample-Aware Guarding Engine for Robust Intrusion Detection Against Adversarial Attacks
by: Chen, Jing, et al.
Published: (2025)
by: Chen, Jing, et al.
Published: (2025)
GaussMarker: Robust Dual-Domain Watermark for Diffusion Models
by: Li, Kecen, et al.
Published: (2025)
by: Li, Kecen, et al.
Published: (2025)
TwoHamsters: Benchmarking Multi-Concept Compositional Unsafety in Text-to-Image Models
by: Zhang, Chaoshuo, et al.
Published: (2026)
by: Zhang, Chaoshuo, et al.
Published: (2026)
SilhouetteTell: Practical Video Identification Leveraging Blurred Recordings of Video Subtitles
by: Huang, Guanchong, et al.
Published: (2025)
by: Huang, Guanchong, et al.
Published: (2025)
Similar Items
-
Do Concept Replacement Techniques Really Erase Unacceptable Concepts?
by: Das, Anudeep, et al.
Published: (2025) -
Bi-Erasing: A Bidirectional Framework for Concept Removal in Diffusion Models
by: Chen, Hao, et al.
Published: (2025) -
Sparse Autoencoder as a Zero-Shot Classifier for Concept Erasing in Text-to-Image Diffusion Models
by: Tian, Zhihua, et al.
Published: (2025) -
Towards Understanding Unsafe Video Generation
by: Pang, Yan, et al.
Published: (2024) -
Collaborative Vision-Text Representation Optimizing for Open-Vocabulary Segmentation
by: Jiao, Siyu, et al.
Published: (2024)