Guarding the Gate: ConceptGuard Battles Concept-Level Backdoors in Concept Bottleneck Models
Fuente:
arXiv
Saved in:
| Main Authors: | Lai, Songning, Huang, Yu, Yang, Jiayu, Huang, Gaoxiang, Chen, Wenshuo, Yue, Yutao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CAT: Concept-level backdoor ATtacks for Concept Bottleneck Models
by: Lai, Songning, et al.
Published: (2024)
by: Lai, Songning, et al.
Published: (2024)
Learning New Concepts, Remembering the Old: Continual Learning for Multimodal Concept Bottleneck Models
by: Lai, Songning, et al.
Published: (2024)
by: Lai, Songning, et al.
Published: (2024)
Mitigating Spurious Background Bias in Multimedia Recognition with Disentangled Concept Bottlenecks
by: Huang, Gaoxiang, et al.
Published: (2025)
by: Huang, Gaoxiang, et al.
Published: (2025)
Backdooring CLIP through Concept Confusion
by: Hu, Lijie, et al.
Published: (2025)
by: Hu, Lijie, et al.
Published: (2025)
Rethinking Robust Adversarial Concept Erasure in Diffusion Models
by: Yin, Qinghong, et al.
Published: (2025)
by: Yin, Qinghong, et al.
Published: (2025)
Do Concept Replacement Techniques Really Erase Unacceptable Concepts?
by: Das, Anudeep, et al.
Published: (2025)
by: Das, Anudeep, et al.
Published: (2025)
GuardSplat: Efficient and Robust Watermarking for 3D Gaussian Splatting
by: Chen, Zixuan, et al.
Published: (2024)
by: Chen, Zixuan, et al.
Published: (2024)
ConceptGuard: Continual Personalized Text-to-Image Generation with Forgetting and Confusion Mitigation
by: Guo, Zirun, et al.
Published: (2025)
by: Guo, Zirun, et al.
Published: (2025)
Six-CD: Benchmarking Concept Removals for Benign Text-to-image Diffusion Models
by: Ren, Jie, et al.
Published: (2024)
by: Ren, Jie, et al.
Published: (2024)
LoyalDiffusion: A Diffusion Model Guarding Against Data Replication
by: Li, Chenghao, et al.
Published: (2024)
by: Li, Chenghao, et al.
Published: (2024)
Bi-Erasing: A Bidirectional Framework for Concept Removal in Diffusion Models
by: Chen, Hao, et al.
Published: (2025)
by: Chen, Hao, et al.
Published: (2025)
Espresso: Robust Concept Filtering in Text-to-Image Models
by: Das, Anudeep, et al.
Published: (2024)
by: Das, Anudeep, et al.
Published: (2024)
TwoHamsters: Benchmarking Multi-Concept Compositional Unsafety in Text-to-Image Models
by: Zhang, Chaoshuo, et al.
Published: (2026)
by: Zhang, Chaoshuo, et al.
Published: (2026)
Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion Models
by: Zhang, Yimeng, et al.
Published: (2024)
by: Zhang, Yimeng, et al.
Published: (2024)
ConceptGuard: Proactive Safety in Text-and-Image-to-Video Generation through Multimodal Risk Detection
by: Ma, Ruize, et al.
Published: (2025)
by: Ma, Ruize, et al.
Published: (2025)
Neighbor-Aware Localized Concept Erasure in Text-to-Image Diffusion Models
by: Shi, Zhuan, et al.
Published: (2026)
by: Shi, Zhuan, et al.
Published: (2026)
IdentityGuard: Context-Aware Restriction and Provenance for Personalized Synthesis
by: Zhang, Lingyun, et al.
Published: (2026)
by: Zhang, Lingyun, et al.
Published: (2026)
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
by: Yuan, Lingzhi, et al.
Published: (2025)
by: Yuan, Lingzhi, et al.
Published: (2025)
VideoEraser: Concept Erasure in Text-to-Video Diffusion Models
by: Xu, Naen, et al.
Published: (2025)
by: Xu, Naen, et al.
Published: (2025)
CGCE: Classifier-Guided Concept Erasure in Generative Models
by: Nguyen, Viet, et al.
Published: (2025)
by: Nguyen, Viet, et al.
Published: (2025)
SAGE: Exploring the Boundaries of Unsafe Concept Domain with Semantic-Augment Erasing
by: Zhu, Hongguang, et al.
Published: (2025)
by: Zhu, Hongguang, et al.
Published: (2025)
Detecting Malicious Concepts without Image Generation in AI-Generated Content (AIGC)
by: Xu, Kun, et al.
Published: (2025)
by: Xu, Kun, et al.
Published: (2025)
What Concepts Lie Within? Detecting and Suppressing Risky Content in Diffusion Transformers
by: Zhang, Chenyu
Published: (2026)
by: Zhang, Chenyu
Published: (2026)
Beyond Text Prompts: Precise Concept Erasure through Text-Image Collaboration
by: Li, Jun, et al.
Published: (2026)
by: Li, Jun, et al.
Published: (2026)
FreezeAsGuard: Mitigating Illegal Adaptation of Diffusion Models via Selective Tensor Freezing
by: Huang, Kai, et al.
Published: (2024)
by: Huang, Kai, et al.
Published: (2024)
VLM-Guard: Safeguarding Vision-Language Models via Fulfilling Safety Alignment Gap
by: Liu, Qin, et al.
Published: (2025)
by: Liu, Qin, et al.
Published: (2025)
Unstable Unlearning: The Hidden Risk of Concept Resurgence in Diffusion Models
by: Suriyakumar, Vinith M., et al.
Published: (2024)
by: Suriyakumar, Vinith M., et al.
Published: (2024)
Blind Spot Navigation: Evolutionary Discovery of Sensitive Semantic Concepts for LVLMs
by: Pan, Zihao, et al.
Published: (2025)
by: Pan, Zihao, et al.
Published: (2025)
GuardTrace-VL: Detecting Unsafe Multimodel Reasoning via Iterative Safety Supervision
by: Xiang, Yuxiao, et al.
Published: (2025)
by: Xiang, Yuxiao, et al.
Published: (2025)
CP-uniGuard: A Unified, Probability-Agnostic, and Adaptive Framework for Malicious Agent Detection and Defense in Multi-Agent Embodied Perception Systems
by: Hu, Senkang, et al.
Published: (2025)
by: Hu, Senkang, et al.
Published: (2025)
IConMark: Robust Interpretable Concept-Based Watermark For AI Images
by: Sadasivan, Vinu Sankar, et al.
Published: (2025)
by: Sadasivan, Vinu Sankar, et al.
Published: (2025)
Sparse Autoencoder as a Zero-Shot Classifier for Concept Erasing in Text-to-Image Diffusion Models
by: Tian, Zhihua, et al.
Published: (2025)
by: Tian, Zhihua, et al.
Published: (2025)
Backdoor Attacks on Prompt-Driven Video Segmentation Foundation Models
by: Zhang, Zongmin, et al.
Published: (2025)
by: Zhang, Zongmin, et al.
Published: (2025)
Backdoor Directions in Vision Transformers
by: Karayalcin, Sengim, et al.
Published: (2026)
by: Karayalcin, Sengim, et al.
Published: (2026)
Meta-Unlearning on Diffusion Models: Preventing Relearning Unlearned Concepts
by: Gao, Hongcheng, et al.
Published: (2024)
by: Gao, Hongcheng, et al.
Published: (2024)
BadBlocks: Low-Cost and Stealthy Backdoor Attacks Tailored for Text-to-Image Diffusion Models
by: Wu, Jia, et al.
Published: (2025)
by: Wu, Jia, et al.
Published: (2025)
Clean-image Backdoor Attacks
by: Rong, Dazhong, et al.
Published: (2024)
by: Rong, Dazhong, et al.
Published: (2024)
Poisoning-based Backdoor Attacks for Arbitrary Target Label with Positive Triggers
by: Huang, Binxiao, et al.
Published: (2024)
by: Huang, Binxiao, et al.
Published: (2024)
What Lurks Within? Concept Auditing for Shared Diffusion Models at Scale
by: Yuan, Xiaoyong, et al.
Published: (2025)
by: Yuan, Xiaoyong, et al.
Published: (2025)
Model X-ray:Detecting Backdoored Models via Decision Boundary
by: Su, Yanghao, et al.
Published: (2024)
by: Su, Yanghao, et al.
Published: (2024)
Similar Items
-
CAT: Concept-level backdoor ATtacks for Concept Bottleneck Models
by: Lai, Songning, et al.
Published: (2024) -
Learning New Concepts, Remembering the Old: Continual Learning for Multimodal Concept Bottleneck Models
by: Lai, Songning, et al.
Published: (2024) -
Mitigating Spurious Background Bias in Multimedia Recognition with Disentangled Concept Bottlenecks
by: Huang, Gaoxiang, et al.
Published: (2025) -
Backdooring CLIP through Concept Confusion
by: Hu, Lijie, et al.
Published: (2025) -
Rethinking Robust Adversarial Concept Erasure in Diffusion Models
by: Yin, Qinghong, et al.
Published: (2025)