Embedding Hidden Adversarial Capabilities in Pre-Trained Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Beerens, Lucas, Higham, Desmond J. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Deceptive Diffusion: Generating Synthetic Adversarial Examples
by: Beerens, Lucas, et al.
Published: (2024)
by: Beerens, Lucas, et al.
Published: (2024)
Rethinking the Vulnerability of Concept Erasure and a New Method
by: Richardson, Alex D., et al.
Published: (2025)
by: Richardson, Alex D., et al.
Published: (2025)
Unveiling the Backdoor Mechanism Hidden Behind Catastrophic Overfitting in Fast Adversarial Training
by: Zhao, Mengnan, et al.
Published: (2026)
by: Zhao, Mengnan, et al.
Published: (2026)
Adversarial Illusions in Multi-Modal Embeddings
by: Zhang, Tingwei, et al.
Published: (2023)
by: Zhang, Tingwei, et al.
Published: (2023)
Information Theoretic Adversarial Training of Large Language Models
by: Zhang, Yiwei, et al.
Published: (2026)
by: Zhang, Yiwei, et al.
Published: (2026)
Generating Adversarial Point Clouds Using Diffusion Model
by: Zhao, Ruiyang, et al.
Published: (2025)
by: Zhao, Ruiyang, et al.
Published: (2025)
Untargeted Adversarial Attack on Knowledge Graph Embeddings
by: Zhao, Tianzhe, et al.
Published: (2024)
by: Zhao, Tianzhe, et al.
Published: (2024)
Exploring Privacy and Fairness Risks in Sharing Diffusion Models: An Adversarial Perspective
by: Luo, Xinjian, et al.
Published: (2024)
by: Luo, Xinjian, et al.
Published: (2024)
Protecting Copyright of Medical Pre-trained Language Models: Training-Free Backdoor Model Watermarking
by: Kong, Cong, et al.
Published: (2024)
by: Kong, Cong, et al.
Published: (2024)
Diffusion-based Adversarial Purification for Intrusion Detection
by: Merzouk, Mohamed Amine, et al.
Published: (2024)
by: Merzouk, Mohamed Amine, et al.
Published: (2024)
Disttack: Graph Adversarial Attacks Toward Distributed GNN Training
by: Zhang, Yuxiang, et al.
Published: (2024)
by: Zhang, Yuxiang, et al.
Published: (2024)
Defending Against Unforeseen Failure Modes with Latent Adversarial Training
by: Casper, Stephen, et al.
Published: (2024)
by: Casper, Stephen, et al.
Published: (2024)
LoRID: Low-Rank Iterative Diffusion for Adversarial Purification
by: Zollicoffer, Geigh, et al.
Published: (2024)
by: Zollicoffer, Geigh, et al.
Published: (2024)
TEESlice: Protecting Sensitive Neural Network Models in Trusted Execution Environments When Attackers have Pre-Trained Models
by: Li, Ding, et al.
Published: (2024)
by: Li, Ding, et al.
Published: (2024)
Jailbroken Frontier Models Retain Their Capabilities
by: Zhu, Daniel, et al.
Published: (2026)
by: Zhu, Daniel, et al.
Published: (2026)
AEGIS: Adversarial Target-Guided Retention-Data-Free Robust Concept Erasure from Diffusion Models
by: Li, Fengpeng, et al.
Published: (2026)
by: Li, Fengpeng, et al.
Published: (2026)
NPAT Null-Space Projected Adversarial Training Towards Zero Deterioration
by: Hu, Hanyi, et al.
Published: (2024)
by: Hu, Hanyi, et al.
Published: (2024)
Generalist++: A Meta-learning Framework for Mitigating Trade-off in Adversarial Training
by: Wang, Yisen, et al.
Published: (2025)
by: Wang, Yisen, et al.
Published: (2025)
Improving Clean Accuracy via a Tangent-Space Perspective on Adversarial Training
by: Yi, Bongsoo, et al.
Published: (2024)
by: Yi, Bongsoo, et al.
Published: (2024)
Automated Creation of Source Code Variants of a Cryptographic Hash Function Implementation Using Generative Pre-Trained Transformer Models
by: Pelofske, Elijah, et al.
Published: (2024)
by: Pelofske, Elijah, et al.
Published: (2024)
Adversaries Can Misuse Combinations of Safe Models
by: Jones, Erik, et al.
Published: (2024)
by: Jones, Erik, et al.
Published: (2024)
PPT-GNN: A Practical Pre-Trained Spatio-Temporal Graph Neural Network for Network Security
by: Van Langendonck, Louis, et al.
Published: (2024)
by: Van Langendonck, Louis, et al.
Published: (2024)
PHANTOM: Progressive High-fidelity Adversarial Network for Threat Object Modeling
by: Al-Karaki, Jamal, et al.
Published: (2025)
by: Al-Karaki, Jamal, et al.
Published: (2025)
Modeling Behavioral Preferences of Cyber Adversaries Using Inverse Reinforcement Learning
by: Shinde, Aditya, et al.
Published: (2025)
by: Shinde, Aditya, et al.
Published: (2025)
MF-CLIP: Leveraging CLIP as Surrogate Models for No-box Adversarial Attacks
by: Zhang, Jiaming, et al.
Published: (2023)
by: Zhang, Jiaming, et al.
Published: (2023)
Adversarial Distilled Retrieval-Augmented Guarding Model for Online Malicious Intent Detection
by: Guo, Yihao, et al.
Published: (2025)
by: Guo, Yihao, et al.
Published: (2025)
Amnesia: Adversarial Semantic Layer Specific Activation Steering in Large Language Models
by: Raza, Ali, et al.
Published: (2026)
by: Raza, Ali, et al.
Published: (2026)
Watermarking Diffusion Language Models
by: Gloaguen, Thibaud, et al.
Published: (2025)
by: Gloaguen, Thibaud, et al.
Published: (2025)
Vision Transformer with Adversarial Indicator Token against Adversarial Attacks in Radio Signal Classifications
by: Zhang, Lu, et al.
Published: (2025)
by: Zhang, Lu, et al.
Published: (2025)
One Token Embedding Is Enough to Deadlock Your Large Reasoning Model
by: Zhang, Mohan, et al.
Published: (2025)
by: Zhang, Mohan, et al.
Published: (2025)
Improved Large Language Model Jailbreak Detection via Pretrained Embeddings
by: Galinkin, Erick, et al.
Published: (2024)
by: Galinkin, Erick, et al.
Published: (2024)
An AI Architecture with the Capability to Classify and Explain Hardware Trojans
by: Whitten, Paul, et al.
Published: (2024)
by: Whitten, Paul, et al.
Published: (2024)
Stateless Yet Not Forgetful: Implicit Memory as a Hidden Channel in LLMs
by: Salem, Ahmed, et al.
Published: (2026)
by: Salem, Ahmed, et al.
Published: (2026)
Self-interpreting Adversarial Images
by: Zhang, Tingwei, et al.
Published: (2024)
by: Zhang, Tingwei, et al.
Published: (2024)
GenoArmory: A Unified Evaluation Framework for Adversarial Attacks on Genomic Foundation Models
by: Luo, Haozheng, et al.
Published: (2025)
by: Luo, Haozheng, et al.
Published: (2025)
BEACON: Behavioral Malware Classification with Large Language Model Embeddings and Deep Learning
by: Perera, Wadduwage Shanika, et al.
Published: (2025)
by: Perera, Wadduwage Shanika, et al.
Published: (2025)
Predictive Auditing of Hidden Tokens in LLM APIs via Reasoning Length Estimation
by: Wang, Ziyao, et al.
Published: (2025)
by: Wang, Ziyao, et al.
Published: (2025)
Topological Signatures of Adversaries in Multimodal Alignments
by: Vu, Minh, et al.
Published: (2025)
by: Vu, Minh, et al.
Published: (2025)
Are Robust LLM Fingerprints Adversarially Robust?
by: Nasery, Anshul, et al.
Published: (2025)
by: Nasery, Anshul, et al.
Published: (2025)
Investigating the Impact of Quantization on Adversarial Robustness
by: Li, Qun, et al.
Published: (2024)
by: Li, Qun, et al.
Published: (2024)
Similar Items
-
Deceptive Diffusion: Generating Synthetic Adversarial Examples
by: Beerens, Lucas, et al.
Published: (2024) -
Rethinking the Vulnerability of Concept Erasure and a New Method
by: Richardson, Alex D., et al.
Published: (2025) -
Unveiling the Backdoor Mechanism Hidden Behind Catastrophic Overfitting in Fast Adversarial Training
by: Zhao, Mengnan, et al.
Published: (2026) -
Adversarial Illusions in Multi-Modal Embeddings
by: Zhang, Tingwei, et al.
Published: (2023) -
Information Theoretic Adversarial Training of Large Language Models
by: Zhang, Yiwei, et al.
Published: (2026)