Attention Shift: Steering AI Away from Unsafe Content
Fuente:
arXiv
Guardado en:
| Autores principales: | Garg, Shivank, Tiwari, Manyana |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Unmasking the Veil: An Investigation into Concept Ablation for Privacy and Copyright Protection in Images
por: Garg, Shivank, et al.
Publicado: (2024)
por: Garg, Shivank, et al.
Publicado: (2024)
Mitigating Sexual Content Generation via Embedding Distortion in Text-conditioned Diffusion Models
por: Ahn, Jaesin, et al.
Publicado: (2025)
por: Ahn, Jaesin, et al.
Publicado: (2025)
AICAttack: Adversarial Image Captioning Attack with Attention-Based Optimization
por: Li, Jiyao, et al.
Publicado: (2024)
por: Li, Jiyao, et al.
Publicado: (2024)
Learning Counterfactually Decoupled Attention for Open-World Model Attribution
por: Zheng, Yu, et al.
Publicado: (2025)
por: Zheng, Yu, et al.
Publicado: (2025)
Structure Disruption: Subverting Malicious Diffusion-Based Inpainting via Self-Attention Query Perturbation
por: He, Yuhao, et al.
Publicado: (2025)
por: He, Yuhao, et al.
Publicado: (2025)
Mirage in the Eyes: Hallucination Attack on Multi-modal Large Language Models with Only Attention Sink
por: Wang, Yining, et al.
Publicado: (2025)
por: Wang, Yining, et al.
Publicado: (2025)
Perturbing Attention Gives You More Bang for the Buck: Subtle Imaging Perturbations That Efficiently Fool Customized Diffusion Models
por: Xu, Jingyao, et al.
Publicado: (2024)
por: Xu, Jingyao, et al.
Publicado: (2024)
AI-generated Image Detection: Passive or Watermark?
por: Guo, Moyang, et al.
Publicado: (2024)
por: Guo, Moyang, et al.
Publicado: (2024)
Copyright Protection in Generative AI: A Technical Perspective
por: Ren, Jie, et al.
Publicado: (2024)
por: Ren, Jie, et al.
Publicado: (2024)
Transparency Attacks: How Imperceptible Image Layers Can Fool AI Perception
por: McKee, Forrest, et al.
Publicado: (2024)
por: McKee, Forrest, et al.
Publicado: (2024)
Provenance of AI-Generated Images: A Vector Similarity and Blockchain-based Approach
por: Sharma, Jitendra, et al.
Publicado: (2025)
por: Sharma, Jitendra, et al.
Publicado: (2025)
Detection of AI Deepfake and Fraud in Online Payments Using GAN-Based Models
por: Ke, Zong, et al.
Publicado: (2025)
por: Ke, Zong, et al.
Publicado: (2025)
Watermark-based Attribution of AI-Generated Content
por: Jiang, Zhengyuan, et al.
Publicado: (2024)
por: Jiang, Zhengyuan, et al.
Publicado: (2024)
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
por: Yuan, Lingzhi, et al.
Publicado: (2025)
por: Yuan, Lingzhi, et al.
Publicado: (2025)
Set You Straight: Auto-Steering Denoising Trajectories to Sidestep Unwanted Concepts
por: Li, Leyang, et al.
Publicado: (2025)
por: Li, Leyang, et al.
Publicado: (2025)
RAW: A Robust and Agile Plug-and-Play Watermark Framework for AI-Generated Images with Provable Guarantees
por: Xian, Xun, et al.
Publicado: (2024)
por: Xian, Xun, et al.
Publicado: (2024)
Can a Teenager Fool an AI? Evaluating Low-Cost Cosmetic Attacks on Age Estimation Systems
por: Shen, Xingyu, et al.
Publicado: (2026)
por: Shen, Xingyu, et al.
Publicado: (2026)
DeepfakeArt Challenge: A Benchmark Dataset for Generative AI Art Forgery and Data Poisoning Detection
por: Aboutalebi, Hossein, et al.
Publicado: (2023)
por: Aboutalebi, Hossein, et al.
Publicado: (2023)
Extracting Training Data from Unconditional Diffusion Models
por: Chen, Yunhao, et al.
Publicado: (2024)
por: Chen, Yunhao, et al.
Publicado: (2024)
©Plug-in Authorization for Human Content Copyright Protection in Text-to-Image Model
por: Zhou, Chao, et al.
Publicado: (2024)
por: Zhou, Chao, et al.
Publicado: (2024)
SIDE: Surrogate Conditional Data Extraction from Diffusion Models
por: Chen, Yunhao, et al.
Publicado: (2024)
por: Chen, Yunhao, et al.
Publicado: (2024)
Generating Potent Poisons and Backdoors from Scratch with Guided Diffusion
por: Souri, Hossein, et al.
Publicado: (2024)
por: Souri, Hossein, et al.
Publicado: (2024)
Perturb and Recover: Fine-tuning for Effective Backdoor Removal from CLIP
por: Singh, Naman Deep, et al.
Publicado: (2024)
por: Singh, Naman Deep, et al.
Publicado: (2024)
Exploring Adversarial Attacks against Latent Diffusion Model from the Perspective of Adversarial Transferability
por: Chen, Junxi, et al.
Publicado: (2024)
por: Chen, Junxi, et al.
Publicado: (2024)
Deep-TEMPEST: Using Deep Learning to Eavesdrop on HDMI from its Unintended Electromagnetic Emanations
por: Fernández, Santiago, et al.
Publicado: (2024)
por: Fernández, Santiago, et al.
Publicado: (2024)
Accurate and Private Diagnosis of Rare Genetic Syndromes from Facial Images with Federated Deep Learning
por: Ünal, Ali Burak, et al.
Publicado: (2025)
por: Ünal, Ali Burak, et al.
Publicado: (2025)
Poisoning Attacks on Federated Learning for Autonomous Driving
por: Garg, Sonakshi, et al.
Publicado: (2024)
por: Garg, Sonakshi, et al.
Publicado: (2024)
Towards Privacy-Guaranteed Label Unlearning in Vertical Federated Learning: Few-Shot Forgetting without Disclosure
por: Gu, Hanlin, et al.
Publicado: (2024)
por: Gu, Hanlin, et al.
Publicado: (2024)
SoK: Systematization and Benchmarking of Deepfake Detectors in a Unified Framework
por: Le, Binh M., et al.
Publicado: (2024)
por: Le, Binh M., et al.
Publicado: (2024)
Omni-IML: Towards Unified Image Manipulation Localization
por: Qu, Chenfan, et al.
Publicado: (2024)
por: Qu, Chenfan, et al.
Publicado: (2024)
SnatchML: Hijacking ML models without Training Access
por: Ghorbel, Mahmoud, et al.
Publicado: (2024)
por: Ghorbel, Mahmoud, et al.
Publicado: (2024)
Let the Noise Speak: Harnessing Noise for a Unified Defense Against Adversarial and Backdoor Attacks
por: Shahriar, Md Hasan, et al.
Publicado: (2024)
por: Shahriar, Md Hasan, et al.
Publicado: (2024)
Privacy-Preserving Video Anomaly Detection: A Survey
por: Liu, Yang, et al.
Publicado: (2024)
por: Liu, Yang, et al.
Publicado: (2024)
FedAli: Personalized Federated Learning Alignment with Prototype Layers for Generalized Mobile Services
por: Ek, Sannara, et al.
Publicado: (2024)
por: Ek, Sannara, et al.
Publicado: (2024)
MeanSparse: Post-Training Robustness Enhancement Through Mean-Centered Feature Sparsification
por: Amini, Sajjad, et al.
Publicado: (2024)
por: Amini, Sajjad, et al.
Publicado: (2024)
Attack and Reset for Unlearning: Exploiting Adversarial Noise toward Machine Unlearning through Parameter Re-initialization
por: Jung, Yoonhwa, et al.
Publicado: (2024)
por: Jung, Yoonhwa, et al.
Publicado: (2024)
Towards Accurate and Robust Architectures via Neural Architecture Search
por: Ou, Yuwei, et al.
Publicado: (2024)
por: Ou, Yuwei, et al.
Publicado: (2024)
Time Series Anomaly Detection with CNN for Environmental Sensors in Healthcare-IoT
por: Khatun, Mirza Akhi, et al.
Publicado: (2024)
por: Khatun, Mirza Akhi, et al.
Publicado: (2024)
1-D CNN-Based Online Signature Verification with Federated Learning
por: Zhang, Lingfeng, et al.
Publicado: (2024)
por: Zhang, Lingfeng, et al.
Publicado: (2024)
Position Paper: Think Globally, React Locally -- Bringing Real-time Reference-based Website Phishing Detection on macOS
por: Petrukha, Ivan, et al.
Publicado: (2024)
por: Petrukha, Ivan, et al.
Publicado: (2024)
Ejemplares similares
-
Unmasking the Veil: An Investigation into Concept Ablation for Privacy and Copyright Protection in Images
por: Garg, Shivank, et al.
Publicado: (2024) -
Mitigating Sexual Content Generation via Embedding Distortion in Text-conditioned Diffusion Models
por: Ahn, Jaesin, et al.
Publicado: (2025) -
AICAttack: Adversarial Image Captioning Attack with Attention-Based Optimization
por: Li, Jiyao, et al.
Publicado: (2024) -
Learning Counterfactually Decoupled Attention for Open-World Model Attribution
por: Zheng, Yu, et al.
Publicado: (2025) -
Structure Disruption: Subverting Malicious Diffusion-Based Inpainting via Self-Attention Query Perturbation
por: He, Yuhao, et al.
Publicado: (2025)