Rethinking the Vulnerability of Concept Erasure and a New Method
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Richardson, Alex D., Zhang, Kaicheng, Beerens, Lucas, Chen, Dongdong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Embedding Hidden Adversarial Capabilities in Pre-Trained Diffusion Models
von: Beerens, Lucas, et al.
Veröffentlicht: (2025)
von: Beerens, Lucas, et al.
Veröffentlicht: (2025)
Erased but Not Forgotten: How Backdoors Compromise Concept Erasure
von: Braun, Tobias, et al.
Veröffentlicht: (2025)
von: Braun, Tobias, et al.
Veröffentlicht: (2025)
AEGIS: Adversarial Target-Guided Retention-Data-Free Robust Concept Erasure from Diffusion Models
von: Li, Fengpeng, et al.
Veröffentlicht: (2026)
von: Li, Fengpeng, et al.
Veröffentlicht: (2026)
The Erasure Illusion: Stress-Testing the Generalization of LLM Forgetting Evaluation
von: Jia, Hengrui, et al.
Veröffentlicht: (2025)
von: Jia, Hengrui, et al.
Veröffentlicht: (2025)
Adversarial Vulnerability Under Temporal Concept Drift: A Longitudinal Study of Android Malware Detection
von: Sabbah, Ahmed, et al.
Veröffentlicht: (2026)
von: Sabbah, Ahmed, et al.
Veröffentlicht: (2026)
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
von: Peng, Benji, et al.
Veröffentlicht: (2024)
von: Peng, Benji, et al.
Veröffentlicht: (2024)
Can Neural Decompilation Assist Vulnerability Prediction on Binary Code?
von: Cotroneo, D., et al.
Veröffentlicht: (2024)
von: Cotroneo, D., et al.
Veröffentlicht: (2024)
Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
von: Li, Yuxi, et al.
Veröffentlicht: (2024)
von: Li, Yuxi, et al.
Veröffentlicht: (2024)
Statement-Level Vulnerability Detection: Learning Vulnerability Patterns Through Information Theory and Contrastive Learning
von: Nguyen, Van, et al.
Veröffentlicht: (2022)
von: Nguyen, Van, et al.
Veröffentlicht: (2022)
Learnability and Privacy Vulnerability are Entangled in a Few Critical Weights
von: Fang, Xingli, et al.
Veröffentlicht: (2026)
von: Fang, Xingli, et al.
Veröffentlicht: (2026)
Finetuning Large Language Models for Vulnerability Detection
von: Shestov, Alexey, et al.
Veröffentlicht: (2024)
von: Shestov, Alexey, et al.
Veröffentlicht: (2024)
Rethinking Pruning for Backdoor Mitigation: An Optimization Perspective
von: Li, Nan, et al.
Veröffentlicht: (2024)
von: Li, Nan, et al.
Veröffentlicht: (2024)
Enhancing Vulnerability Reports with Automated and Augmented Description Summarization
von: Althebeiti, Hattan, et al.
Veröffentlicht: (2025)
von: Althebeiti, Hattan, et al.
Veröffentlicht: (2025)
Exploiting Efficiency Vulnerabilities in Dynamic Deep Learning Systems
von: Rathnasuriya, Ravishka, et al.
Veröffentlicht: (2025)
von: Rathnasuriya, Ravishka, et al.
Veröffentlicht: (2025)
ARVO: Atlas of Reproducible Vulnerabilities for Open Source Software
von: Mei, Xiang, et al.
Veröffentlicht: (2024)
von: Mei, Xiang, et al.
Veröffentlicht: (2024)
Syntax- and Compilation-Preserving Evasion of LLM Vulnerability Detectors
von: Sun, Luze, et al.
Veröffentlicht: (2026)
von: Sun, Luze, et al.
Veröffentlicht: (2026)
Large Language Models in Cybersecurity: Applications, Vulnerabilities, and Defense Techniques
von: Jaffal, Niveen O., et al.
Veröffentlicht: (2025)
von: Jaffal, Niveen O., et al.
Veröffentlicht: (2025)
Efficient but Vulnerable: Benchmarking and Defending LLM Batch Prompting Attack
von: Yue, Murong, et al.
Veröffentlicht: (2025)
von: Yue, Murong, et al.
Veröffentlicht: (2025)
Weakest Link in the Chain: Security Vulnerabilities in Advanced Reasoning Models
von: Krishna, Arjun, et al.
Veröffentlicht: (2025)
von: Krishna, Arjun, et al.
Veröffentlicht: (2025)
Securing Large Language Models: Threats, Vulnerabilities and Responsible Practices
von: Abdali, Sara, et al.
Veröffentlicht: (2024)
von: Abdali, Sara, et al.
Veröffentlicht: (2024)
Exploiting Layer-Specific Vulnerabilities to Backdoor Attack in Federated Learning
von: Foroughi, Mohammad Hadi, et al.
Veröffentlicht: (2026)
von: Foroughi, Mohammad Hadi, et al.
Veröffentlicht: (2026)
An Unbiased Transformer Source Code Learning with Semantic Vulnerability Graph
von: Islam, Nafis Tanveer, et al.
Veröffentlicht: (2023)
von: Islam, Nafis Tanveer, et al.
Veröffentlicht: (2023)
Your Privacy Depends on Others: Collusion Vulnerabilities in Individual Differential Privacy
von: Kaiser, Johannes, et al.
Veröffentlicht: (2026)
von: Kaiser, Johannes, et al.
Veröffentlicht: (2026)
Detection of False Data Injection Attacks (FDIA) on Power Dynamical Systems With a State Prediction Method
von: Sahu, Abhijeet, et al.
Veröffentlicht: (2024)
von: Sahu, Abhijeet, et al.
Veröffentlicht: (2024)
Probing Network Decisions: Capturing Uncertainties and Unveiling Vulnerabilities Without Label Information
von: Joung, Youngju, et al.
Veröffentlicht: (2025)
von: Joung, Youngju, et al.
Veröffentlicht: (2025)
ChatBug: A Common Vulnerability of Aligned LLMs Induced by Chat Templates
von: Jiang, Fengqing, et al.
Veröffentlicht: (2024)
von: Jiang, Fengqing, et al.
Veröffentlicht: (2024)
When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models
von: Hossain, Ismail, et al.
Veröffentlicht: (2026)
von: Hossain, Ismail, et al.
Veröffentlicht: (2026)
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
von: Wang, Zhun, et al.
Veröffentlicht: (2026)
von: Wang, Zhun, et al.
Veröffentlicht: (2026)
Injection, Attack and Erasure: Revocable Backdoor Attacks via Machine Unlearning
von: Song, Baogang, et al.
Veröffentlicht: (2025)
von: Song, Baogang, et al.
Veröffentlicht: (2025)
AED: Automatic Discovery of Effective and Diverse Vulnerabilities for Autonomous Driving Policy with Large Language Models
von: Qiu, Le, et al.
Veröffentlicht: (2025)
von: Qiu, Le, et al.
Veröffentlicht: (2025)
BugSweeper: Function-Level Detection of Smart Contract Vulnerabilities Using Graph Neural Networks
von: Lee, Uisang, et al.
Veröffentlicht: (2025)
von: Lee, Uisang, et al.
Veröffentlicht: (2025)
Empirical Evaluation of Concept Drift in ML-Based Android Malware Detection
von: Sabbah, Ahmed, et al.
Veröffentlicht: (2025)
von: Sabbah, Ahmed, et al.
Veröffentlicht: (2025)
Optimized Deep Learning Models for Malware Detection under Concept Drift
von: Maillet, William, et al.
Veröffentlicht: (2023)
von: Maillet, William, et al.
Veröffentlicht: (2023)
SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert Identification
von: Lai, Zhenglin, et al.
Veröffentlicht: (2025)
von: Lai, Zhenglin, et al.
Veröffentlicht: (2025)
How to make Medical AI Systems safer? Simulating Vulnerabilities, and Threats in Multimodal Medical RAG System
von: Zuo, Kaiwen, et al.
Veröffentlicht: (2025)
von: Zuo, Kaiwen, et al.
Veröffentlicht: (2025)
Logic layer Prompt Control Injection (LPCI): A Novel Security Vulnerability Class in Agentic Systems
von: Atta, Hammad, et al.
Veröffentlicht: (2025)
von: Atta, Hammad, et al.
Veröffentlicht: (2025)
Dynamic Neural Control Flow Execution: An Agent-Based Deep Equilibrium Approach for Binary Vulnerability Detection
von: Li, Litao, et al.
Veröffentlicht: (2024)
von: Li, Litao, et al.
Veröffentlicht: (2024)
Concept Drift Adaptation Using Self-Supervised and Reinforcement Learning In Android Malware Detection
von: Sabbah, Ahmed, et al.
Veröffentlicht: (2026)
von: Sabbah, Ahmed, et al.
Veröffentlicht: (2026)
What's Pulling the Strings? Evaluating Integrity and Attribution in AI Training and Inference through Concept Shift
von: Chang, Jiamin, et al.
Veröffentlicht: (2025)
von: Chang, Jiamin, et al.
Veröffentlicht: (2025)
Deep Learning-Driven Malware Classification with API Call Sequence Analysis and Concept Drift Handling
von: Gond, Bishwajit Prasad, et al.
Veröffentlicht: (2025)
von: Gond, Bishwajit Prasad, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Embedding Hidden Adversarial Capabilities in Pre-Trained Diffusion Models
von: Beerens, Lucas, et al.
Veröffentlicht: (2025) -
Erased but Not Forgotten: How Backdoors Compromise Concept Erasure
von: Braun, Tobias, et al.
Veröffentlicht: (2025) -
AEGIS: Adversarial Target-Guided Retention-Data-Free Robust Concept Erasure from Diffusion Models
von: Li, Fengpeng, et al.
Veröffentlicht: (2026) -
The Erasure Illusion: Stress-Testing the Generalization of LLM Forgetting Evaluation
von: Jia, Hengrui, et al.
Veröffentlicht: (2025) -
Adversarial Vulnerability Under Temporal Concept Drift: A Longitudinal Study of Android Malware Detection
von: Sabbah, Ahmed, et al.
Veröffentlicht: (2026)