Model X-ray:Detecting Backdoored Models via Decision Boundary
Fuente:
arXiv
Saved in:
| Main Authors: | Su, Yanghao, Zhang, Jie, Xu, Ting, Zhang, Tianwei, Zhang, Weiming, Yu, Nenghai |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BURN: Backdoor Unlearning via Adversarial Boundary Analysis
by: Su, Yanghao, et al.
Published: (2025)
by: Su, Yanghao, et al.
Published: (2025)
Gaussian Shading: Provable Performance-Lossless Image Watermarking for Diffusion Models
by: Yang, Zijin, et al.
Published: (2024)
by: Yang, Zijin, et al.
Published: (2024)
SafeGuider: Robust and Practical Content Safety Control for Text-to-Image Models
by: Qi, Peigui, et al.
Published: (2025)
by: Qi, Peigui, et al.
Published: (2025)
Gaussian Shading++: Rethinking the Realistic Deployment Challenge of Performance-Lossless Image Watermark for Diffusion Models
by: Yang, Zijin, et al.
Published: (2025)
by: Yang, Zijin, et al.
Published: (2025)
Natias: Neuron Attribution based Transferable Image Adversarial Steganography
by: Fan, Zexin, et al.
Published: (2024)
by: Fan, Zexin, et al.
Published: (2024)
Model Supply Chain Poisoning: Backdooring Pre-trained Models via Embedding Indistinguishability
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
Revocable Backdoor for Deep Model Trading
by: Xu, Yiran, et al.
Published: (2024)
by: Xu, Yiran, et al.
Published: (2024)
©Plug-in Authorization for Human Content Copyright Protection in Text-to-Image Model
by: Zhou, Chao, et al.
Published: (2024)
by: Zhou, Chao, et al.
Published: (2024)
LiteUpdate: A Lightweight Framework for Updating AI-Generated Image Detectors
by: Lu, Jiajie, et al.
Published: (2025)
by: Lu, Jiajie, et al.
Published: (2025)
Character as a Latent Variable in Large Language Models: A Mechanistic Account of Emergent Misalignment and Conditional Safety Failures
by: Su, Yanghao, et al.
Published: (2026)
by: Su, Yanghao, et al.
Published: (2026)
SemBind: Binding Diffusion Watermarks to Semantics Against Black-Box Forgery Attacks
by: Zhang, Xin, et al.
Published: (2026)
by: Zhang, Xin, et al.
Published: (2026)
Backdoor Attacks against Image-to-Image Networks
by: Jiang, Wenbo, et al.
Published: (2024)
by: Jiang, Wenbo, et al.
Published: (2024)
Gungnir: Exploiting Stylistic Features in Images for Backdoor Attacks on Diffusion Models
by: Zhang, Lei, et al.
Published: (2025)
by: Zhang, Lei, et al.
Published: (2025)
Test-Time Attention Purification for Backdoored Large Vision Language Models
by: Zhang, Zhifang, et al.
Published: (2026)
by: Zhang, Zhifang, et al.
Published: (2026)
Backdoor Attack with Mode Mixture Latent Modification
by: Zhang, Hongwei, et al.
Published: (2024)
by: Zhang, Hongwei, et al.
Published: (2024)
Backdoor Attacks on Prompt-Driven Video Segmentation Foundation Models
by: Zhang, Zongmin, et al.
Published: (2025)
by: Zhang, Zongmin, et al.
Published: (2025)
UNIT: Backdoor Mitigation via Automated Neural Distribution Tightening
by: Cheng, Siyuan, et al.
Published: (2024)
by: Cheng, Siyuan, et al.
Published: (2024)
CatchBackdoor: Backdoor Detection via Critical Trojan Neural Path Fuzzing
by: Jin, Haibo, et al.
Published: (2021)
by: Jin, Haibo, et al.
Published: (2021)
BadBone: Backdoor Attacks Against Backbone Models in Visual Prompt Learning
by: Yang, Ziqing, et al.
Published: (2026)
by: Yang, Ziqing, et al.
Published: (2026)
A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations
by: Ye, Mang, et al.
Published: (2025)
by: Ye, Mang, et al.
Published: (2025)
Evidence-based Decision Modeling for Synthetic Face Detection with Uncertainty-driven Active Learning
by: Jiang, Qingchao, et al.
Published: (2026)
by: Jiang, Qingchao, et al.
Published: (2026)
Backdoor Mitigation in Object Detection via Adversarial Fine-Tuning
by: Dunnett, Kealan, et al.
Published: (2026)
by: Dunnett, Kealan, et al.
Published: (2026)
Backdoor Attacks against No-Reference Image Quality Assessment Models via a Scalable Trigger
by: Yu, Yi, et al.
Published: (2024)
by: Yu, Yi, et al.
Published: (2024)
CIS-BA: Continuous Interaction Space Based Backdoor Attack for Object Detection in the Real-World
by: Zhao, Shuxin, et al.
Published: (2025)
by: Zhao, Shuxin, et al.
Published: (2025)
DisDet: Exploring Detectability of Backdoor Attack on Diffusion Models
by: Sui, Yang, et al.
Published: (2024)
by: Sui, Yang, et al.
Published: (2024)
Scaling Exposes the Trigger: Input-Level Backdoor Detection in Text-to-Image Diffusion Models via Cross-Attention Scaling
by: Li, Zida, et al.
Published: (2026)
by: Li, Zida, et al.
Published: (2026)
LOTUS: Evasive and Resilient Backdoor Attacks through Sub-Partitioning
by: Cheng, Siyuan, et al.
Published: (2024)
by: Cheng, Siyuan, et al.
Published: (2024)
InverTune: Removing Backdoors from Multimodal Contrastive Learning Models via Trigger Inversion and Activation Tuning
by: Sun, Mengyuan, et al.
Published: (2025)
by: Sun, Mengyuan, et al.
Published: (2025)
Goal-oriented Backdoor Attack against Vision-Language-Action Models via Physical Objects
by: Zhou, Zirun, et al.
Published: (2025)
by: Zhou, Zirun, et al.
Published: (2025)
Model Pairing Using Embedding Translation for Backdoor Attack Detection on Open-Set Classification Tasks
by: Unnervik, Alexander, et al.
Published: (2024)
by: Unnervik, Alexander, et al.
Published: (2024)
ConSeg: Contextual Backdoor Attack Against Semantic Segmentation
by: Abbasi, Bilal Hussain, et al.
Published: (2025)
by: Abbasi, Bilal Hussain, et al.
Published: (2025)
HoneypotNet: Backdoor Attacks Against Model Extraction
by: Wang, Yixu, et al.
Published: (2025)
by: Wang, Yixu, et al.
Published: (2025)
Gaussian Shannon: High-Precision Diffusion Model Watermarking Based on Communication
by: Zhang, Yi, et al.
Published: (2026)
by: Zhang, Yi, et al.
Published: (2026)
Guarding the Gate: ConceptGuard Battles Concept-Level Backdoors in Concept Bottleneck Models
by: Lai, Songning, et al.
Published: (2024)
by: Lai, Songning, et al.
Published: (2024)
BadBlocks: Low-Cost and Stealthy Backdoor Attacks Tailored for Text-to-Image Diffusion Models
by: Wu, Jia, et al.
Published: (2025)
by: Wu, Jia, et al.
Published: (2025)
Disrupting Vision-Language Model-Driven Navigation Services via Adversarial Object Fusion
by: Xie, Chunlong, et al.
Published: (2025)
by: Xie, Chunlong, et al.
Published: (2025)
Backdoor Attack on Vision Language Models with Stealthy Semantic Manipulation
by: Zhong, Zhiyuan, et al.
Published: (2025)
by: Zhong, Zhiyuan, et al.
Published: (2025)
Mitigating Backdoor Attacks using Activation-Guided Model Editing
by: Hsieh, Felix, et al.
Published: (2024)
by: Hsieh, Felix, et al.
Published: (2024)
BadDet+: Robust Backdoor Attacks for Object Detection
by: Dunnett, Kealan, et al.
Published: (2026)
by: Dunnett, Kealan, et al.
Published: (2026)
SAME: Sample Reconstruction against Model Extraction Attacks
by: Xie, Yi, et al.
Published: (2023)
by: Xie, Yi, et al.
Published: (2023)
Similar Items
-
BURN: Backdoor Unlearning via Adversarial Boundary Analysis
by: Su, Yanghao, et al.
Published: (2025) -
Gaussian Shading: Provable Performance-Lossless Image Watermarking for Diffusion Models
by: Yang, Zijin, et al.
Published: (2024) -
SafeGuider: Robust and Practical Content Safety Control for Text-to-Image Models
by: Qi, Peigui, et al.
Published: (2025) -
Gaussian Shading++: Rethinking the Realistic Deployment Challenge of Performance-Lossless Image Watermark for Diffusion Models
by: Yang, Zijin, et al.
Published: (2025) -
Natias: Neuron Attribution based Transferable Image Adversarial Steganography
by: Fan, Zexin, et al.
Published: (2024)