Post-train Black-box Defense via Bayesian Boundary Correction
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, He, Diao, Yunfeng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Bounding-box Watermarking: Defense against Model Extraction Attacks on Object Detectors
por: Koda, Satoru, et al.
Publicado: (2024)
por: Koda, Satoru, et al.
Publicado: (2024)
Delving into Decision-based Black-box Attacks on Semantic Segmentation
por: Chen, Zhaoyu, et al.
Publicado: (2024)
por: Chen, Zhaoyu, et al.
Publicado: (2024)
Vulnerabilities in AI-generated Image Detection: The Challenge of Adversarial Attacks
por: Diao, Yunfeng, et al.
Publicado: (2024)
por: Diao, Yunfeng, et al.
Publicado: (2024)
CSF: Black-box Fingerprinting via Compositional Semantics for Text-to-Image Models
por: Lee, Junhoo, et al.
Publicado: (2026)
por: Lee, Junhoo, et al.
Publicado: (2026)
Efficient Black-box Adversarial Attacks via Bayesian Optimization Guided by a Function Prior
por: Cheng, Shuyu, et al.
Publicado: (2024)
por: Cheng, Shuyu, et al.
Publicado: (2024)
Struggle with Adversarial Defense? Try Diffusion
por: Li, Yujie, et al.
Publicado: (2024)
por: Li, Yujie, et al.
Publicado: (2024)
PuriDefense: Randomized Local Implicit Adversarial Purification for Defending Black-box Query-based Attacks
por: Guo, Ping, et al.
Publicado: (2024)
por: Guo, Ping, et al.
Publicado: (2024)
BSPA: Exploring Black-box Stealthy Prompt Attacks against Image Generators
por: Tian, Yu, et al.
Publicado: (2024)
por: Tian, Yu, et al.
Publicado: (2024)
Stego Battlefield: Evaluating Image Steganography Attacks and Steganalysis Defenses
por: Sun, Zhen, et al.
Publicado: (2026)
por: Sun, Zhen, et al.
Publicado: (2026)
PromptLA: Towards Integrity Verification of Black-box Text-to-Image Diffusion Models
por: Zhang, Zhuomeng, et al.
Publicado: (2024)
por: Zhang, Zhuomeng, et al.
Publicado: (2024)
SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore Mechanism
por: Chen, Beitao, et al.
Publicado: (2025)
por: Chen, Beitao, et al.
Publicado: (2025)
MMCert: Provable Defense against Adversarial Attacks to Multi-modal Models
por: Wang, Yanting, et al.
Publicado: (2024)
por: Wang, Yanting, et al.
Publicado: (2024)
Model X-ray:Detecting Backdoored Models via Decision Boundary
por: Su, Yanghao, et al.
Publicado: (2024)
por: Su, Yanghao, et al.
Publicado: (2024)
MIBench: A Comprehensive Framework for Benchmarking Model Inversion Attack and Defense
por: Qiu, Yixiang, et al.
Publicado: (2024)
por: Qiu, Yixiang, et al.
Publicado: (2024)
VLATTACK: Multimodal Adversarial Attacks on Vision-Language Tasks via Pre-trained Models
por: Yin, Ziyi, et al.
Publicado: (2023)
por: Yin, Ziyi, et al.
Publicado: (2023)
Mitigating Backdoor Attack by Injecting Proactive Defensive Backdoor
por: Wei, Shaokui, et al.
Publicado: (2024)
por: Wei, Shaokui, et al.
Publicado: (2024)
Data-free Defense of Black Box Models Against Adversarial Attacks
por: Nayak, Gaurav Kumar, et al.
Publicado: (2022)
por: Nayak, Gaurav Kumar, et al.
Publicado: (2022)
Superpixel Attack: Enhancing Black-box Adversarial Attack with Image-driven Division Areas
por: Oe, Issa, et al.
Publicado: (2025)
por: Oe, Issa, et al.
Publicado: (2025)
A Survey of Trojan Attacks and Defenses to Deep Neural Networks
por: Jin, Lingxin, et al.
Publicado: (2024)
por: Jin, Lingxin, et al.
Publicado: (2024)
The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense
por: Guo, Yangyang, et al.
Publicado: (2024)
por: Guo, Yangyang, et al.
Publicado: (2024)
Revisiting Adversarial Patch Defenses on Object Detectors: Unified Evaluation, Large-Scale Dataset, and New Insights
por: Zheng, Junhao, et al.
Publicado: (2025)
por: Zheng, Junhao, et al.
Publicado: (2025)
Token-Level Constraint Boundary Search for Jailbreaking Text-to-Image Models
por: Liu, Jiangtao, et al.
Publicado: (2025)
por: Liu, Jiangtao, et al.
Publicado: (2025)
Adversarial Defenses via Vector Quantization
por: Dong, Zhiyi, et al.
Publicado: (2023)
por: Dong, Zhiyi, et al.
Publicado: (2023)
Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion Models
por: Zhang, Yimeng, et al.
Publicado: (2024)
por: Zhang, Yimeng, et al.
Publicado: (2024)
Class-feature Watermark: A Resilient Black-box Watermark Against Model Extraction Attacks
por: Xiao, Yaxin, et al.
Publicado: (2025)
por: Xiao, Yaxin, et al.
Publicado: (2025)
ODDR: Outlier Detection & Dimension Reduction Based Defense Against Adversarial Patches
por: Chattopadhyay, Nandish, et al.
Publicado: (2023)
por: Chattopadhyay, Nandish, et al.
Publicado: (2023)
Robust Image Classification: Defensive Strategies against FGSM and PGD Adversarial Attacks
por: Waghela, Hetvi, et al.
Publicado: (2024)
por: Waghela, Hetvi, et al.
Publicado: (2024)
A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations
por: Ye, Mang, et al.
Publicado: (2025)
por: Ye, Mang, et al.
Publicado: (2025)
Nearest is Not Dearest: Towards Practical Defense against Quantization-conditioned Backdoor Attacks
por: Li, Boheng, et al.
Publicado: (2024)
por: Li, Boheng, et al.
Publicado: (2024)
SAP-DIFF: Semantic Adversarial Patch Generation for Black-Box Face Recognition Models via Diffusion Models
por: Wang, Mingsi, et al.
Publicado: (2025)
por: Wang, Mingsi, et al.
Publicado: (2025)
DeMark: A Query-Free Black-Box Attack on Deepfake Watermarking Defenses
por: Song, Wei, et al.
Publicado: (2026)
por: Song, Wei, et al.
Publicado: (2026)
Hot-Swap MarkBoard: An Efficient Black-box Watermarking Approach for Large-scale Model Distribution
por: Zhang, Zhicheng, et al.
Publicado: (2025)
por: Zhang, Zhicheng, et al.
Publicado: (2025)
Defensive Adversarial CAPTCHA: A Semantics-Driven Framework for Natural Adversarial Example Generation
por: Du, Xia, et al.
Publicado: (2025)
por: Du, Xia, et al.
Publicado: (2025)
PatchCURE: Improving Certifiable Robustness, Model Utility, and Computation Efficiency of Adversarial Patch Defenses
por: Xiang, Chong, et al.
Publicado: (2023)
por: Xiang, Chong, et al.
Publicado: (2023)
Training-Free Color-Aware Adversarial Diffusion Sanitization for Diffusion Stegomalware Defense at Security Gateways
por: Frants, Vladimir, et al.
Publicado: (2025)
por: Frants, Vladimir, et al.
Publicado: (2025)
Continual Adversarial Defense
por: Wang, Qian, et al.
Publicado: (2023)
por: Wang, Qian, et al.
Publicado: (2023)
OmniSafeBench-MM: A Unified Benchmark and Toolbox for Multimodal Jailbreak Attack-Defense Evaluation
por: Jia, Xiaojun, et al.
Publicado: (2025)
por: Jia, Xiaojun, et al.
Publicado: (2025)
Model Supply Chain Poisoning: Backdooring Pre-trained Models via Embedding Indistinguishability
por: Wang, Hao, et al.
Publicado: (2024)
por: Wang, Hao, et al.
Publicado: (2024)
Hijack-GAN: Unintended-Use of Pretrained, Black-Box GANs
por: Wang, Hui-Po, et al.
Publicado: (2020)
por: Wang, Hui-Po, et al.
Publicado: (2020)
Backdoor Defense in Diffusion Models via Spatial Attention Unlearning
por: Jha, Abha, et al.
Publicado: (2025)
por: Jha, Abha, et al.
Publicado: (2025)
Ejemplares similares
-
Bounding-box Watermarking: Defense against Model Extraction Attacks on Object Detectors
por: Koda, Satoru, et al.
Publicado: (2024) -
Delving into Decision-based Black-box Attacks on Semantic Segmentation
por: Chen, Zhaoyu, et al.
Publicado: (2024) -
Vulnerabilities in AI-generated Image Detection: The Challenge of Adversarial Attacks
por: Diao, Yunfeng, et al.
Publicado: (2024) -
CSF: Black-box Fingerprinting via Compositional Semantics for Text-to-Image Models
por: Lee, Junhoo, et al.
Publicado: (2026) -
Efficient Black-box Adversarial Attacks via Bayesian Optimization Guided by a Function Prior
por: Cheng, Shuyu, et al.
Publicado: (2024)