Confidence-aware Denoised Fine-tuning of Off-the-shelf Models for Certified Robustness
Fuente:
arXiv
Saved in:
| Main Authors: | Jang, Suhyeok, Kim, Seojin, Shin, Jinwoo, Jeong, Jongheon |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
StarFT: Robust Fine-tuning of Zero-shot Models via Spuriosity Alignment
by: Kim, Younghyun, et al.
Published: (2025)
by: Kim, Younghyun, et al.
Published: (2025)
Autoencoder-based Denoising Defense against Adversarial Attacks on Object Detection
by: Song, Min Geun, et al.
Published: (2025)
by: Song, Min Geun, et al.
Published: (2025)
Towards Resilient Safety-driven Unlearning for Diffusion Models against Downstream Fine-tuning
by: Li, Boheng, et al.
Published: (2025)
by: Li, Boheng, et al.
Published: (2025)
Differentially Private Fine-Tuning of Diffusion Models
by: Tsai, Yu-Lin, et al.
Published: (2024)
by: Tsai, Yu-Lin, et al.
Published: (2024)
BackWeak: Backdooring Knowledge Distillation Simply with Weak Triggers and Fine-tuning
by: Wang, Shanmin, et al.
Published: (2025)
by: Wang, Shanmin, et al.
Published: (2025)
Multi-View Slot Attention Using Paraphrased Texts for Face Anti-Spoofing
by: Yu, Jeongmin, et al.
Published: (2025)
by: Yu, Jeongmin, et al.
Published: (2025)
Adversarial Robustness of Vision in Open Foundation Models
by: Fox, Jonathon, et al.
Published: (2025)
by: Fox, Jonathon, et al.
Published: (2025)
Silent Branding Attack: Trigger-free Data Poisoning Attack on Text-to-Image Diffusion Models
by: Jang, Sangwon, et al.
Published: (2025)
by: Jang, Sangwon, et al.
Published: (2025)
ROBIN: Robust and Invisible Watermarks for Diffusion Models with Adversarial Optimization
by: Huang, Huayang, et al.
Published: (2024)
by: Huang, Huayang, et al.
Published: (2024)
Certified Robustness to Clean-Label Poisoning Using Diffusion Denoising
by: Hong, Sanghyun, et al.
Published: (2024)
by: Hong, Sanghyun, et al.
Published: (2024)
Enhancing Variational Autoencoders with Smooth Robust Latent Encoding
by: Lee, Hyomin, et al.
Published: (2025)
by: Lee, Hyomin, et al.
Published: (2025)
On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression
by: Zhang, Xinwei, et al.
Published: (2026)
by: Zhang, Xinwei, et al.
Published: (2026)
Evaluating the Efficacy of Prompt-Engineered Large Multimodal Models Versus Fine-Tuned Vision Transformers in Image-Based Security Applications
by: Trad, Fouad, et al.
Published: (2024)
by: Trad, Fouad, et al.
Published: (2024)
GIFT: Gradient-aware Immunization of diffusion models against malicious Fine-Tuning with safe concepts retention
by: Abdalla, Amro, et al.
Published: (2025)
by: Abdalla, Amro, et al.
Published: (2025)
Watertox: The Art of Simplicity in Universal Attacks A Cross-Model Framework for Robust Adversarial Generation
by: Gao, Zhenghao, et al.
Published: (2024)
by: Gao, Zhenghao, et al.
Published: (2024)
Unveiling Hidden Visual Information: A Reconstruction Attack Against Adversarial Visual Information Hiding
by: Jang, Jonggyu, et al.
Published: (2024)
by: Jang, Jonggyu, et al.
Published: (2024)
Can VLMs Detect and Localize Fine-Grained AI-Edited Images?
by: Sun, Zhen, et al.
Published: (2025)
by: Sun, Zhen, et al.
Published: (2025)
On the Robustness of Watermarking for Autoregressive Image Generation
by: Müller, Andreas, et al.
Published: (2026)
by: Müller, Andreas, et al.
Published: (2026)
PAD-FT: A Lightweight Defense for Backdoor Attacks via Data Purification and Fine-Tuning
by: Xu, Yukai, et al.
Published: (2024)
by: Xu, Yukai, et al.
Published: (2024)
Effective Fine-Tuning of Vision Transformers with Low-Rank Adaptation for Privacy-Preserving Image Classification
by: Lin, Haiwei, et al.
Published: (2025)
by: Lin, Haiwei, et al.
Published: (2025)
Cert-SSBD: Certified Backdoor Defense with Sample-Specific Smoothing Noises
by: Qiao, Ting, et al.
Published: (2025)
by: Qiao, Ting, et al.
Published: (2025)
Adaptive Diffusion Denoised Smoothing : Certified Robustness via Randomized Smoothing with Differentially Private Guided Denoising Diffusion
by: Shpilevskiy, Frederick, et al.
Published: (2025)
by: Shpilevskiy, Frederick, et al.
Published: (2025)
CertDW: Towards Certified Dataset Ownership Verification via Conformal Prediction
by: Qiao, Ting, et al.
Published: (2025)
by: Qiao, Ting, et al.
Published: (2025)
RoboCurate: Harnessing Diversity with Action-Verified Neural Trajectory for Robot Learning
by: Kim, Seungku, et al.
Published: (2026)
by: Kim, Seungku, et al.
Published: (2026)
SafeVision: Efficient Image Guardrail with Robust Policy Adherence and Explainability
by: Xu, Peiyang, et al.
Published: (2025)
by: Xu, Peiyang, et al.
Published: (2025)
Robustness Analysis against Adversarial Patch Attacks in Fully Unmanned Stores
by: Na, Hyunsik, et al.
Published: (2025)
by: Na, Hyunsik, et al.
Published: (2025)
IConMark: Robust Interpretable Concept-Based Watermark For AI Images
by: Sadasivan, Vinu Sankar, et al.
Published: (2025)
by: Sadasivan, Vinu Sankar, et al.
Published: (2025)
Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM Security
by: Fang, Xiang, et al.
Published: (2026)
by: Fang, Xiang, et al.
Published: (2026)
Robust Watermarking Using Generative Priors Against Image Editing: From Benchmarking to Advances
by: Lu, Shilin, et al.
Published: (2024)
by: Lu, Shilin, et al.
Published: (2024)
Naïve Exposure of Generative AI Capabilities Undermines Deepfake Detection
by: Kim, Sunpill, et al.
Published: (2026)
by: Kim, Sunpill, et al.
Published: (2026)
Rel-Zero: Harnessing Patch-Pair Invariance for Robust Zero-Watermarking Against AI Editing
by: Chen, Pengzhen, et al.
Published: (2026)
by: Chen, Pengzhen, et al.
Published: (2026)
Set You Straight: Auto-Steering Denoising Trajectories to Sidestep Unwanted Concepts
by: Li, Leyang, et al.
Published: (2025)
by: Li, Leyang, et al.
Published: (2025)
CtrlAttack: A Unified Attack on World-Model Control in Diffusion Models
by: Xu, Shuhan, et al.
Published: (2026)
by: Xu, Shuhan, et al.
Published: (2026)
Gradient Inversion of Federated Diffusion Models
by: Huang, Jiyue, et al.
Published: (2024)
by: Huang, Jiyue, et al.
Published: (2024)
PromptSmooth: Certifying Robustness of Medical Vision-Language Models via Prompt Learning
by: Hussein, Noor, et al.
Published: (2024)
by: Hussein, Noor, et al.
Published: (2024)
Learning to Watermark in the Latent Space of Generative Models
by: Rebuffi, Sylvestre-Alvise, et al.
Published: (2026)
by: Rebuffi, Sylvestre-Alvise, et al.
Published: (2026)
Region-Guided Attack on the Segment Anything Model (SAM)
by: Liu, Xiaoliang, et al.
Published: (2024)
by: Liu, Xiaoliang, et al.
Published: (2024)
CGCE: Classifier-Guided Concept Erasure in Generative Models
by: Nguyen, Viet, et al.
Published: (2025)
by: Nguyen, Viet, et al.
Published: (2025)
Metaphor-based Jailbreak Attacks on Text-to-Image Models
by: Zhang, Chenyu, et al.
Published: (2025)
by: Zhang, Chenyu, et al.
Published: (2025)
Rethinking and Red-Teaming Protective Perturbation in Personalized Diffusion Models
by: Liu, Yixin, et al.
Published: (2024)
by: Liu, Yixin, et al.
Published: (2024)
Similar Items
-
StarFT: Robust Fine-tuning of Zero-shot Models via Spuriosity Alignment
by: Kim, Younghyun, et al.
Published: (2025) -
Autoencoder-based Denoising Defense against Adversarial Attacks on Object Detection
by: Song, Min Geun, et al.
Published: (2025) -
Towards Resilient Safety-driven Unlearning for Diffusion Models against Downstream Fine-tuning
by: Li, Boheng, et al.
Published: (2025) -
Differentially Private Fine-Tuning of Diffusion Models
by: Tsai, Yu-Lin, et al.
Published: (2024) -
BackWeak: Backdooring Knowledge Distillation Simply with Weak Triggers and Fine-tuning
by: Wang, Shanmin, et al.
Published: (2025)