Do Concept Replacement Techniques Really Erase Unacceptable Concepts?
Fuente:
arXiv
Saved in:
| Main Authors: | Das, Anudeep, Singh, Gurjot, Chantasantitam, Prach, Asokan, N. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Espresso: Robust Concept Filtering in Text-to-Image Models
by: Das, Anudeep, et al.
Published: (2024)
by: Das, Anudeep, et al.
Published: (2024)
Backdooring Bias in Large Language Models
by: Das, Anudeep, et al.
Published: (2026)
by: Das, Anudeep, et al.
Published: (2026)
ReVision : A Post-Hoc, Vision-Based Technique for Replacing Unacceptable Concepts in Image Generation Pipeline
by: Singh, Gurjot, et al.
Published: (2026)
by: Singh, Gurjot, et al.
Published: (2026)
Bi-Erasing: A Bidirectional Framework for Concept Removal in Diffusion Models
by: Chen, Hao, et al.
Published: (2025)
by: Chen, Hao, et al.
Published: (2025)
SAGE: Exploring the Boundaries of Unsafe Concept Domain with Semantic-Augment Erasing
by: Zhu, Hongguang, et al.
Published: (2025)
by: Zhu, Hongguang, et al.
Published: (2025)
Sparse Autoencoder as a Zero-Shot Classifier for Concept Erasing in Text-to-Image Diffusion Models
by: Tian, Zhihua, et al.
Published: (2025)
by: Tian, Zhihua, et al.
Published: (2025)
Guarding the Gate: ConceptGuard Battles Concept-Level Backdoors in Concept Bottleneck Models
by: Lai, Songning, et al.
Published: (2024)
by: Lai, Songning, et al.
Published: (2024)
CAT: Concept-level backdoor ATtacks for Concept Bottleneck Models
by: Lai, Songning, et al.
Published: (2024)
by: Lai, Songning, et al.
Published: (2024)
Backdooring CLIP through Concept Confusion
by: Hu, Lijie, et al.
Published: (2025)
by: Hu, Lijie, et al.
Published: (2025)
PAL*M: Property Attestation for Large Generative Models
by: Chantasantitam, Prach, et al.
Published: (2026)
by: Chantasantitam, Prach, et al.
Published: (2026)
Rethinking Robust Adversarial Concept Erasure in Diffusion Models
by: Yin, Qinghong, et al.
Published: (2025)
by: Yin, Qinghong, et al.
Published: (2025)
Neighbor-Aware Localized Concept Erasure in Text-to-Image Diffusion Models
by: Shi, Zhuan, et al.
Published: (2026)
by: Shi, Zhuan, et al.
Published: (2026)
Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion Models
by: Zhang, Yimeng, et al.
Published: (2024)
by: Zhang, Yimeng, et al.
Published: (2024)
Detecting Malicious Concepts without Image Generation in AI-Generated Content (AIGC)
by: Xu, Kun, et al.
Published: (2025)
by: Xu, Kun, et al.
Published: (2025)
What Concepts Lie Within? Detecting and Suppressing Risky Content in Diffusion Transformers
by: Zhang, Chenyu
Published: (2026)
by: Zhang, Chenyu
Published: (2026)
Beyond Text Prompts: Precise Concept Erasure through Text-Image Collaboration
by: Li, Jun, et al.
Published: (2026)
by: Li, Jun, et al.
Published: (2026)
Six-CD: Benchmarking Concept Removals for Benign Text-to-image Diffusion Models
by: Ren, Jie, et al.
Published: (2024)
by: Ren, Jie, et al.
Published: (2024)
TwoHamsters: Benchmarking Multi-Concept Compositional Unsafety in Text-to-Image Models
by: Zhang, Chaoshuo, et al.
Published: (2026)
by: Zhang, Chaoshuo, et al.
Published: (2026)
FedSECA: Sign Election and Coordinate-wise Aggregation of Gradients for Byzantine Tolerant Federated Learning
by: Benjamin, Joseph Geo, et al.
Published: (2024)
by: Benjamin, Joseph Geo, et al.
Published: (2024)
Learning New Concepts, Remembering the Old: Continual Learning for Multimodal Concept Bottleneck Models
by: Lai, Songning, et al.
Published: (2024)
by: Lai, Songning, et al.
Published: (2024)
CGCE: Classifier-Guided Concept Erasure in Generative Models
by: Nguyen, Viet, et al.
Published: (2025)
by: Nguyen, Viet, et al.
Published: (2025)
BlindU: Blind Machine Unlearning without Revealing Erasing Data
by: Wang, Weiqi, et al.
Published: (2026)
by: Wang, Weiqi, et al.
Published: (2026)
VideoEraser: Concept Erasure in Text-to-Video Diffusion Models
by: Xu, Naen, et al.
Published: (2025)
by: Xu, Naen, et al.
Published: (2025)
Unstable Unlearning: The Hidden Risk of Concept Resurgence in Diffusion Models
by: Suriyakumar, Vinith M., et al.
Published: (2024)
by: Suriyakumar, Vinith M., et al.
Published: (2024)
Blind Spot Navigation: Evolutionary Discovery of Sensitive Semantic Concepts for LVLMs
by: Pan, Zihao, et al.
Published: (2025)
by: Pan, Zihao, et al.
Published: (2025)
IConMark: Robust Interpretable Concept-Based Watermark For AI Images
by: Sadasivan, Vinu Sankar, et al.
Published: (2025)
by: Sadasivan, Vinu Sankar, et al.
Published: (2025)
Random Erasing vs. Model Inversion: A Promising Defense or a False Hope?
by: Tran, Viet-Hung, et al.
Published: (2024)
by: Tran, Viet-Hung, et al.
Published: (2024)
SoK: Can Synthetic Images Replace Real Data? A Survey of Utility and Privacy of Synthetic Image Generation
by: Chung, Yunsung, et al.
Published: (2025)
by: Chung, Yunsung, et al.
Published: (2025)
Meta-Unlearning on Diffusion Models: Preventing Relearning Unlearned Concepts
by: Gao, Hongcheng, et al.
Published: (2024)
by: Gao, Hongcheng, et al.
Published: (2024)
A Privacy-Preserving Semantic-Segmentation Method Using Domain-Adaptation Technique
by: Sueyoshi, Homare, et al.
Published: (2025)
by: Sueyoshi, Homare, et al.
Published: (2025)
BEACON: A Multimodal Dataset for Learning Behavioral Fingerprints from Gameplay Data
by: Singh, Ishpuneet, et al.
Published: (2026)
by: Singh, Ishpuneet, et al.
Published: (2026)
DLOVE: A new Security Evaluation Tool for Deep Learning Based Watermarking Techniques
by: Padhi, Sudev Kumar, et al.
Published: (2024)
by: Padhi, Sudev Kumar, et al.
Published: (2024)
Do Modern Post-Hoc Watermarking Methods Beat Broken-Arrows?
by: Gesny, Enoal, et al.
Published: (2026)
by: Gesny, Enoal, et al.
Published: (2026)
Erasing Self-Supervised Learning Backdoor by Cluster Activation Masking
by: Qian, Shengsheng, et al.
Published: (2023)
by: Qian, Shengsheng, et al.
Published: (2023)
Image-to-Text Logic Jailbreak: Your Imagination can Help You Do Anything
by: Zou, Xiaotian, et al.
Published: (2024)
by: Zou, Xiaotian, et al.
Published: (2024)
Do Not Leave a Gap: Hallucination-Free Object Concealment in Vision-Language Models
by: Guesmi, Amira, et al.
Published: (2026)
by: Guesmi, Amira, et al.
Published: (2026)
DeiTFake: Deepfake Detection Model using DeiT Multi-Stage Training
by: Kumar, Saksham, et al.
Published: (2025)
by: Kumar, Saksham, et al.
Published: (2025)
Transformer-Driven Multimodal Fusion for Explainable Suspiciousness Estimation in Visual Surveillance
by: Yadav, Kuldeep Singh, et al.
Published: (2025)
by: Yadav, Kuldeep Singh, et al.
Published: (2025)
What Lurks Within? Concept Auditing for Shared Diffusion Models at Scale
by: Yuan, Xiaoyong, et al.
Published: (2025)
by: Yuan, Xiaoyong, et al.
Published: (2025)
CoreMark: Toward Robust and Universal Text Watermarking Technique
by: Meng, Jiale, et al.
Published: (2025)
by: Meng, Jiale, et al.
Published: (2025)
Similar Items
-
Espresso: Robust Concept Filtering in Text-to-Image Models
by: Das, Anudeep, et al.
Published: (2024) -
Backdooring Bias in Large Language Models
by: Das, Anudeep, et al.
Published: (2026) -
ReVision : A Post-Hoc, Vision-Based Technique for Replacing Unacceptable Concepts in Image Generation Pipeline
by: Singh, Gurjot, et al.
Published: (2026) -
Bi-Erasing: A Bidirectional Framework for Concept Removal in Diffusion Models
by: Chen, Hao, et al.
Published: (2025) -
SAGE: Exploring the Boundaries of Unsafe Concept Domain with Semantic-Augment Erasing
by: Zhu, Hongguang, et al.
Published: (2025)