CAT: Concept-level backdoor ATtacks for Concept Bottleneck Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lai, Songning, Yang, Jiayu, Huang, Yu, Hu, Lijie, Xue, Tianlang, Hu, Zhangyi, Li, Jiaxu, Liao, Haicheng, Yue, Yutao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning New Concepts, Remembering the Old: Continual Learning for Multimodal Concept Bottleneck Models
von: Lai, Songning, et al.
Veröffentlicht: (2024)
von: Lai, Songning, et al.
Veröffentlicht: (2024)
Guarding the Gate: ConceptGuard Battles Concept-Level Backdoors in Concept Bottleneck Models
von: Lai, Songning, et al.
Veröffentlicht: (2024)
von: Lai, Songning, et al.
Veröffentlicht: (2024)
Backdooring CLIP through Concept Confusion
von: Hu, Lijie, et al.
Veröffentlicht: (2025)
von: Hu, Lijie, et al.
Veröffentlicht: (2025)
Mitigating Spurious Background Bias in Multimedia Recognition with Disentangled Concept Bottlenecks
von: Huang, Gaoxiang, et al.
Veröffentlicht: (2025)
von: Huang, Gaoxiang, et al.
Veröffentlicht: (2025)
Rethinking Robust Adversarial Concept Erasure in Diffusion Models
von: Yin, Qinghong, et al.
Veröffentlicht: (2025)
von: Yin, Qinghong, et al.
Veröffentlicht: (2025)
Do Concept Replacement Techniques Really Erase Unacceptable Concepts?
von: Das, Anudeep, et al.
Veröffentlicht: (2025)
von: Das, Anudeep, et al.
Veröffentlicht: (2025)
Six-CD: Benchmarking Concept Removals for Benign Text-to-image Diffusion Models
von: Ren, Jie, et al.
Veröffentlicht: (2024)
von: Ren, Jie, et al.
Veröffentlicht: (2024)
DRIVE: Dependable Robust Interpretable Visionary Ensemble Framework in Autonomous Driving
von: Lai, Songning, et al.
Veröffentlicht: (2024)
von: Lai, Songning, et al.
Veröffentlicht: (2024)
SlowBA: An efficiency backdoor attack towards VLM-based GUI agents
von: Li, Junxian, et al.
Veröffentlicht: (2026)
von: Li, Junxian, et al.
Veröffentlicht: (2026)
TwoHamsters: Benchmarking Multi-Concept Compositional Unsafety in Text-to-Image Models
von: Zhang, Chaoshuo, et al.
Veröffentlicht: (2026)
von: Zhang, Chaoshuo, et al.
Veröffentlicht: (2026)
Espresso: Robust Concept Filtering in Text-to-Image Models
von: Das, Anudeep, et al.
Veröffentlicht: (2024)
von: Das, Anudeep, et al.
Veröffentlicht: (2024)
NBA: defensive distillation for backdoor removal via neural behavior alignment
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
From base cases to backdoors: An Empirical Study of Unnatural Crypto-API Misuse
von: Olaiya, Victor, et al.
Veröffentlicht: (2025)
von: Olaiya, Victor, et al.
Veröffentlicht: (2025)
DLP: towards active defense against backdoor attacks with decoupled learning process
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
Neighbor-Aware Localized Concept Erasure in Text-to-Image Diffusion Models
von: Shi, Zhuan, et al.
Veröffentlicht: (2026)
von: Shi, Zhuan, et al.
Veröffentlicht: (2026)
Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion Models
von: Zhang, Yimeng, et al.
Veröffentlicht: (2024)
von: Zhang, Yimeng, et al.
Veröffentlicht: (2024)
Bi-Erasing: A Bidirectional Framework for Concept Removal in Diffusion Models
von: Chen, Hao, et al.
Veröffentlicht: (2025)
von: Chen, Hao, et al.
Veröffentlicht: (2025)
Meta-Unlearning on Diffusion Models: Preventing Relearning Unlearned Concepts
von: Gao, Hongcheng, et al.
Veröffentlicht: (2024)
von: Gao, Hongcheng, et al.
Veröffentlicht: (2024)
Detecting Malicious Concepts without Image Generation in AI-Generated Content (AIGC)
von: Xu, Kun, et al.
Veröffentlicht: (2025)
von: Xu, Kun, et al.
Veröffentlicht: (2025)
What Concepts Lie Within? Detecting and Suppressing Risky Content in Diffusion Transformers
von: Zhang, Chenyu
Veröffentlicht: (2026)
von: Zhang, Chenyu
Veröffentlicht: (2026)
Beyond Text Prompts: Precise Concept Erasure through Text-Image Collaboration
von: Li, Jun, et al.
Veröffentlicht: (2026)
von: Li, Jun, et al.
Veröffentlicht: (2026)
A clean-label graph backdoor attack method in node classification task
von: Xing, Xiaogang, et al.
Veröffentlicht: (2023)
von: Xing, Xiaogang, et al.
Veröffentlicht: (2023)
On the critical path to implant backdoors and the effectiveness of potential mitigation techniques: Early learnings from XZ
von: Lins, Mario, et al.
Veröffentlicht: (2024)
von: Lins, Mario, et al.
Veröffentlicht: (2024)
SAGE: Exploring the Boundaries of Unsafe Concept Domain with Semantic-Augment Erasing
von: Zhu, Hongguang, et al.
Veröffentlicht: (2025)
von: Zhu, Hongguang, et al.
Veröffentlicht: (2025)
Model-agnostic clean-label backdoor mitigation in cybersecurity environments
von: Severi, Giorgio, et al.
Veröffentlicht: (2024)
von: Severi, Giorgio, et al.
Veröffentlicht: (2024)
SAP-DIFF: Semantic Adversarial Patch Generation for Black-Box Face Recognition Models via Diffusion Models
von: Wang, Mingsi, et al.
Veröffentlicht: (2025)
von: Wang, Mingsi, et al.
Veröffentlicht: (2025)
Blind Spot Navigation: Evolutionary Discovery of Sensitive Semantic Concepts for LVLMs
von: Pan, Zihao, et al.
Veröffentlicht: (2025)
von: Pan, Zihao, et al.
Veröffentlicht: (2025)
CGCE: Classifier-Guided Concept Erasure in Generative Models
von: Nguyen, Viet, et al.
Veröffentlicht: (2025)
von: Nguyen, Viet, et al.
Veröffentlicht: (2025)
Key Concepts and Principles of Blockchain Technology
von: Ghorbian, Mohsen, et al.
Veröffentlicht: (2025)
von: Ghorbian, Mohsen, et al.
Veröffentlicht: (2025)
Keystroke Dynamics: Concepts, Techniques, and Applications
von: Shadman, Rashik, et al.
Veröffentlicht: (2023)
von: Shadman, Rashik, et al.
Veröffentlicht: (2023)
ImpNet: Imperceptible and blackbox-undetectable backdoors in compiled neural networks
von: Clifford, Eleanor, et al.
Veröffentlicht: (2022)
von: Clifford, Eleanor, et al.
Veröffentlicht: (2022)
TarPro: Targeted Protection against Malicious Image Editing
von: Shen, Kaixin, et al.
Veröffentlicht: (2025)
von: Shen, Kaixin, et al.
Veröffentlicht: (2025)
Aligning Core Aspects: Improving Vulnerability Proof-of-Concepts via Cross-Source Insights
von: Wang, Lingxiao, et al.
Veröffentlicht: (2025)
von: Wang, Lingxiao, et al.
Veröffentlicht: (2025)
Unstable Unlearning: The Hidden Risk of Concept Resurgence in Diffusion Models
von: Suriyakumar, Vinith M., et al.
Veröffentlicht: (2024)
von: Suriyakumar, Vinith M., et al.
Veröffentlicht: (2024)
VideoEraser: Concept Erasure in Text-to-Video Diffusion Models
von: Xu, Naen, et al.
Veröffentlicht: (2025)
von: Xu, Naen, et al.
Veröffentlicht: (2025)
JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation
von: Zhang, Shenyi, et al.
Veröffentlicht: (2025)
von: Zhang, Shenyi, et al.
Veröffentlicht: (2025)
A general approach to enhance the survivability of backdoor attacks by decision path coupling
von: Zhao, Yufei, et al.
Veröffentlicht: (2024)
von: Zhao, Yufei, et al.
Veröffentlicht: (2024)
IConMark: Robust Interpretable Concept-Based Watermark For AI Images
von: Sadasivan, Vinu Sankar, et al.
Veröffentlicht: (2025)
von: Sadasivan, Vinu Sankar, et al.
Veröffentlicht: (2025)
Concept-ROT: Poisoning Concepts in Large Language Models with Model Editing
von: Grimes, Keltin, et al.
Veröffentlicht: (2024)
von: Grimes, Keltin, et al.
Veröffentlicht: (2024)
Transcending Transcend: Revisiting Malware Classification in the Presence of Concept Drift
von: Barbero, Federico, et al.
Veröffentlicht: (2020)
von: Barbero, Federico, et al.
Veröffentlicht: (2020)
Ähnliche Einträge
-
Learning New Concepts, Remembering the Old: Continual Learning for Multimodal Concept Bottleneck Models
von: Lai, Songning, et al.
Veröffentlicht: (2024) -
Guarding the Gate: ConceptGuard Battles Concept-Level Backdoors in Concept Bottleneck Models
von: Lai, Songning, et al.
Veröffentlicht: (2024) -
Backdooring CLIP through Concept Confusion
von: Hu, Lijie, et al.
Veröffentlicht: (2025) -
Mitigating Spurious Background Bias in Multimedia Recognition with Disentangled Concept Bottlenecks
von: Huang, Gaoxiang, et al.
Veröffentlicht: (2025) -
Rethinking Robust Adversarial Concept Erasure in Diffusion Models
von: Yin, Qinghong, et al.
Veröffentlicht: (2025)