Set You Straight: Auto-Steering Denoising Trajectories to Sidestep Unwanted Concepts
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Leyang, Lu, Shilin, Ren, Yan, Kong, Adams Wai-Kin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
All That Glitters Is Not Gold: Key-Secured 3D Secrets within 3D Gaussian Splatting
by: Ren, Yan, et al.
Published: (2025)
by: Ren, Yan, et al.
Published: (2025)
Robust Watermarking Using Generative Priors Against Image Editing: From Benchmarking to Advances
by: Lu, Shilin, et al.
Published: (2024)
by: Lu, Shilin, et al.
Published: (2024)
MACE: Mass Concept Erasure in Diffusion Models
by: Lu, Shilin, et al.
Published: (2024)
by: Lu, Shilin, et al.
Published: (2024)
SteerDiff: Steering towards Safe Text-to-Image Diffusion Models
by: Zhang, Hongxiang, et al.
Published: (2024)
by: Zhang, Hongxiang, et al.
Published: (2024)
SAGE: Exploring the Boundaries of Unsafe Concept Domain with Semantic-Augment Erasing
by: Zhu, Hongguang, et al.
Published: (2025)
by: Zhu, Hongguang, et al.
Published: (2025)
Are You Copying My Prompt? Protecting the Copyright of Vision Prompt for VPaaS via Watermark
by: Ren, Huali, et al.
Published: (2024)
by: Ren, Huali, et al.
Published: (2024)
Autoencoder-based Denoising Defense against Adversarial Attacks on Object Detection
by: Song, Min Geun, et al.
Published: (2025)
by: Song, Min Geun, et al.
Published: (2025)
PromptLA: Towards Integrity Verification of Black-box Text-to-Image Diffusion Models
by: Zhang, Zhuomeng, et al.
Published: (2024)
by: Zhang, Zhuomeng, et al.
Published: (2024)
Text is All You Need for Vision-Language Model Jailbreaking
by: Chen, Yihang, et al.
Published: (2026)
by: Chen, Yihang, et al.
Published: (2026)
CGCE: Classifier-Guided Concept Erasure in Generative Models
by: Nguyen, Viet, et al.
Published: (2025)
by: Nguyen, Viet, et al.
Published: (2025)
Can You Trust What You See? Alpha Channel No-Box Attacks on Video Object Detection
by: Yi, Ariana, et al.
Published: (2025)
by: Yi, Ariana, et al.
Published: (2025)
VideoEraser: Concept Erasure in Text-to-Video Diffusion Models
by: Xu, Naen, et al.
Published: (2025)
by: Xu, Naen, et al.
Published: (2025)
FoolSDEdit: Deceptively Steering Your Edits Towards Targeted Attribute-aware Distribution
by: Zhou, Qi, et al.
Published: (2024)
by: Zhou, Qi, et al.
Published: (2024)
Blind Spot Navigation: Evolutionary Discovery of Sensitive Semantic Concepts for LVLMs
by: Pan, Zihao, et al.
Published: (2025)
by: Pan, Zihao, et al.
Published: (2025)
IConMark: Robust Interpretable Concept-Based Watermark For AI Images
by: Sadasivan, Vinu Sankar, et al.
Published: (2025)
by: Sadasivan, Vinu Sankar, et al.
Published: (2025)
Who Can See Through You? Adversarial Shielding Against VLM-Based Attribute Inference Attacks
by: Fan, Yucheng, et al.
Published: (2025)
by: Fan, Yucheng, et al.
Published: (2025)
Sparse Autoencoder as a Zero-Shot Classifier for Concept Erasing in Text-to-Image Diffusion Models
by: Tian, Zhihua, et al.
Published: (2025)
by: Tian, Zhihua, et al.
Published: (2025)
Guarding the Gate: ConceptGuard Battles Concept-Level Backdoors in Concept Bottleneck Models
by: Lai, Songning, et al.
Published: (2024)
by: Lai, Songning, et al.
Published: (2024)
FedPalm: A General Federated Learning Framework for Closed- and Open-Set Palmprint Verification
by: Yang, Ziyuan, et al.
Published: (2025)
by: Yang, Ziyuan, et al.
Published: (2025)
Six-CD: Benchmarking Concept Removals for Benign Text-to-image Diffusion Models
by: Ren, Jie, et al.
Published: (2024)
by: Ren, Jie, et al.
Published: (2024)
CAT: Concept-level backdoor ATtacks for Concept Bottleneck Models
by: Lai, Songning, et al.
Published: (2024)
by: Lai, Songning, et al.
Published: (2024)
Do Concept Replacement Techniques Really Erase Unacceptable Concepts?
by: Das, Anudeep, et al.
Published: (2025)
by: Das, Anudeep, et al.
Published: (2025)
Detecting AutoAttack Perturbations in the Frequency Domain
by: Lorenz, Peter, et al.
Published: (2021)
by: Lorenz, Peter, et al.
Published: (2021)
Is It Really You? Exploring Biometric Verification Scenarios in Photorealistic Talking-Head Avatar Videos
by: Pedrouzo-Rodriguez, Laura, et al.
Published: (2025)
by: Pedrouzo-Rodriguez, Laura, et al.
Published: (2025)
Invisible Watermarks, Visible Gains: Steering Machine Unlearning with Bi-Level Watermarking Design
by: Sun, Yuhao, et al.
Published: (2025)
by: Sun, Yuhao, et al.
Published: (2025)
Confidence-aware Denoised Fine-tuning of Off-the-shelf Models for Certified Robustness
by: Jang, Suhyeok, et al.
Published: (2024)
by: Jang, Suhyeok, et al.
Published: (2024)
Backdooring CLIP through Concept Confusion
by: Hu, Lijie, et al.
Published: (2025)
by: Hu, Lijie, et al.
Published: (2025)
Is RobustBench/AutoAttack a suitable Benchmark for Adversarial Robustness?
by: Lorenz, Peter, et al.
Published: (2021)
by: Lorenz, Peter, et al.
Published: (2021)
Temporal-Guided Spiking Neural Networks for Event-Based Human Action Recognition
by: Yang, Siyuan, et al.
Published: (2025)
by: Yang, Siyuan, et al.
Published: (2025)
Transient Adversarial 3D Projection Attacks on Object Detection in Autonomous Driving
by: Zhou, Ce, et al.
Published: (2024)
by: Zhou, Ce, et al.
Published: (2024)
The Adversarial AI-Art: Understanding, Generation, Detection, and Benchmarking
by: Li, Yuying, et al.
Published: (2024)
by: Li, Yuying, et al.
Published: (2024)
AutoMIA: Improved Baselines for Membership Inference Attack via Agentic Self-Exploration
by: Liu, Ruhao, et al.
Published: (2026)
by: Liu, Ruhao, et al.
Published: (2026)
Rethinking Robust Adversarial Concept Erasure in Diffusion Models
by: Yin, Qinghong, et al.
Published: (2025)
by: Yin, Qinghong, et al.
Published: (2025)
Espresso: Robust Concept Filtering in Text-to-Image Models
by: Das, Anudeep, et al.
Published: (2024)
by: Das, Anudeep, et al.
Published: (2024)
FC-Attack: Jailbreaking Multimodal Large Language Models via Auto-Generated Flowcharts
by: Zhang, Ziyi, et al.
Published: (2025)
by: Zhang, Ziyi, et al.
Published: (2025)
Backdoor Attack Against Vision Transformers via Attention Gradient-Based Image Erosion
by: Guo, Ji, et al.
Published: (2024)
by: Guo, Ji, et al.
Published: (2024)
AnywhereDoor: Multi-Target Backdoor Attacks on Object Detection
by: Lu, Jialin, et al.
Published: (2024)
by: Lu, Jialin, et al.
Published: (2024)
AnywhereDoor: Multi-Target Backdoor Attacks on Object Detection
by: Lu, Jialin, et al.
Published: (2025)
by: Lu, Jialin, et al.
Published: (2025)
Combating Falsification of Speech Videos with Live Optical Signatures (Extended Version)
by: Schwartz, Hadleigh, et al.
Published: (2025)
by: Schwartz, Hadleigh, et al.
Published: (2025)
L-AutoDA: Leveraging Large Language Models for Automated Decision-based Adversarial Attacks
by: Guo, Ping, et al.
Published: (2024)
by: Guo, Ping, et al.
Published: (2024)
Similar Items
-
All That Glitters Is Not Gold: Key-Secured 3D Secrets within 3D Gaussian Splatting
by: Ren, Yan, et al.
Published: (2025) -
Robust Watermarking Using Generative Priors Against Image Editing: From Benchmarking to Advances
by: Lu, Shilin, et al.
Published: (2024) -
MACE: Mass Concept Erasure in Diffusion Models
by: Lu, Shilin, et al.
Published: (2024) -
SteerDiff: Steering towards Safe Text-to-Image Diffusion Models
by: Zhang, Hongxiang, et al.
Published: (2024) -
SAGE: Exploring the Boundaries of Unsafe Concept Domain with Semantic-Augment Erasing
by: Zhu, Hongguang, et al.
Published: (2025)