Harnessing Hyperbolic Geometry for Harmful Prompt Detection and Sanitization
Fuente:
arXiv
Saved in:
| Main Authors: | Maljkovic, Igor, Briglia, Maria Rosaria, Masi, Iacopo, Cinà, Antonio Emanuele, Roli, Fabio |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Robust image classification with multi-modal large language models
by: Villani, Francesco, et al.
Published: (2024)
by: Villani, Francesco, et al.
Published: (2024)
Machine Learning Security against Data Poisoning: Are We There Yet?
by: Cinà, Antonio Emanuele, et al.
Published: (2022)
by: Cinà, Antonio Emanuele, et al.
Published: (2022)
Evaluating the Evaluators: Trust in Adversarial Robustness Tests
by: Cinà, Antonio Emanuele, et al.
Published: (2025)
by: Cinà, Antonio Emanuele, et al.
Published: (2025)
Energy-Latency Attacks via Sponge Poisoning
by: Cinà, Antonio Emanuele, et al.
Published: (2022)
by: Cinà, Antonio Emanuele, et al.
Published: (2022)
Sonic: Fast and Transferable Data Poisoning on Clustering Algorithms
by: Villani, Francesco, et al.
Published: (2024)
by: Villani, Francesco, et al.
Published: (2024)
AttackBench: Evaluating Gradient-based Attacks for Adversarial Examples
by: Cinà, Antonio Emanuele, et al.
Published: (2024)
by: Cinà, Antonio Emanuele, et al.
Published: (2024)
Implicit Inversion turns CLIP into a Decoder
by: D'Orazio, Antonio, et al.
Published: (2025)
by: D'Orazio, Antonio, et al.
Published: (2025)
On the Generation and Mitigation of Harmful Geometry in Image-to-3D Models
by: Liu, Yule, et al.
Published: (2026)
by: Liu, Yule, et al.
Published: (2026)
Proactive Adversarial Defense: Harnessing Prompt Tuning in Vision-Language Models to Detect Unseen Backdoored Images
by: Stein, Kyle, et al.
Published: (2024)
by: Stein, Kyle, et al.
Published: (2024)
$σ$-zero: Gradient-based Optimization of $\ell_0$-norm Adversarial Examples
by: Cinà, Antonio Emanuele, et al.
Published: (2024)
by: Cinà, Antonio Emanuele, et al.
Published: (2024)
Training-Free Color-Aware Adversarial Diffusion Sanitization for Diffusion Stegomalware Defense at Security Gateways
by: Frants, Vladimir, et al.
Published: (2025)
by: Frants, Vladimir, et al.
Published: (2025)
Unveiling the Potential: Harnessing Deep Metric Learning to Circumvent Video Streaming Encryption
by: Gansekoele, Arwin, et al.
Published: (2024)
by: Gansekoele, Arwin, et al.
Published: (2024)
Toward Accountable AI-Generated Content on Social Platforms: Steganographic Attribution and Multimodal Harm Detection
by: Guan, Xinlei, et al.
Published: (2026)
by: Guan, Xinlei, et al.
Published: (2026)
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
by: Yuan, Lingzhi, et al.
Published: (2025)
by: Yuan, Lingzhi, et al.
Published: (2025)
Rel-Zero: Harnessing Patch-Pair Invariance for Robust Zero-Watermarking Against AI Editing
by: Chen, Pengzhen, et al.
Published: (2026)
by: Chen, Pengzhen, et al.
Published: (2026)
Adversarial Pruning: A Survey and Benchmark of Pruning Methods for Adversarial Robustness
by: Piras, Giorgio, et al.
Published: (2024)
by: Piras, Giorgio, et al.
Published: (2024)
Shedding More Light on Robust Classifiers under the lens of Energy-based Models
by: Mirza, Mujtaba Hussain, et al.
Published: (2024)
by: Mirza, Mujtaba Hussain, et al.
Published: (2024)
What is Adversarial Training for Diffusion Models?
by: Rosaria, Briglia Maria, et al.
Published: (2025)
by: Rosaria, Briglia Maria, et al.
Published: (2025)
PLA: Prompt Learning Attack against Text-to-Image Generative Models
by: Lyu, Xinqi, et al.
Published: (2025)
by: Lyu, Xinqi, et al.
Published: (2025)
Fooling the Watchers: Breaking AIGC Detectors via Semantic Prompt Attacks
by: Hao, Run, et al.
Published: (2025)
by: Hao, Run, et al.
Published: (2025)
ImageNet-Patch: A Dataset for Benchmarking Machine Learning Robustness against Adversarial Patches
by: Pintor, Maura, et al.
Published: (2022)
by: Pintor, Maura, et al.
Published: (2022)
Mind the Trojan Horse: Image Prompt Adapter Enabling Scalable and Deceptive Jailbreaking
by: Chen, Junxi, et al.
Published: (2025)
by: Chen, Junxi, et al.
Published: (2025)
Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM Security
by: Fang, Xiang, et al.
Published: (2026)
by: Fang, Xiang, et al.
Published: (2026)
AttackNet: Enhancing Biometric Security via Tailored Convolutional Neural Network Architectures for Liveness Detection
by: Kuznetsov, Oleksandr, et al.
Published: (2024)
by: Kuznetsov, Oleksandr, et al.
Published: (2024)
PromptLA: Towards Integrity Verification of Black-box Text-to-Image Diffusion Models
by: Zhang, Zhuomeng, et al.
Published: (2024)
by: Zhang, Zhuomeng, et al.
Published: (2024)
Cross-Database Liveness Detection: Insights from Comparative Biometric Analysis
by: Kuznetsov, Oleksandr, et al.
Published: (2024)
by: Kuznetsov, Oleksandr, et al.
Published: (2024)
Image-based Prompt Injection: Hijacking Multimodal LLMs through Visually Embedded Adversarial Instructions
by: Nagaraja, Neha, et al.
Published: (2026)
by: Nagaraja, Neha, et al.
Published: (2026)
T2I-RiskyPrompt: A Benchmark for Safety Evaluation, Attack, and Defense on Text-to-Image Model
by: Zhang, Chenyu, et al.
Published: (2025)
by: Zhang, Chenyu, et al.
Published: (2025)
Evaluating the Efficacy of Prompt-Engineered Large Multimodal Models Versus Fine-Tuned Vision Transformers in Image-Based Security Applications
by: Trad, Fouad, et al.
Published: (2024)
by: Trad, Fouad, et al.
Published: (2024)
Combating Digitally Altered Images: Deepfake Detection
by: Kumar, Saksham, et al.
Published: (2025)
by: Kumar, Saksham, et al.
Published: (2025)
CLIP-Inspector: Model-Level Backdoor Detection for Prompt-Tuned CLIP via OOD Trigger Inversion
by: Jindal, Akshit, et al.
Published: (2026)
by: Jindal, Akshit, et al.
Published: (2026)
The Adversarial AI-Art: Understanding, Generation, Detection, and Benchmarking
by: Li, Yuying, et al.
Published: (2024)
by: Li, Yuying, et al.
Published: (2024)
Comparative Evaluation of Deep Learning Models for Fake Image Detection
by: Pakala, Akhitha, et al.
Published: (2026)
by: Pakala, Akhitha, et al.
Published: (2026)
Wukong Framework for Not Safe For Work Detection in Text-to-Image systems
by: Liu, Mingrui, et al.
Published: (2025)
by: Liu, Mingrui, et al.
Published: (2025)
AnywhereDoor: Multi-Target Backdoor Attacks on Object Detection
by: Lu, Jialin, et al.
Published: (2024)
by: Lu, Jialin, et al.
Published: (2024)
AnywhereDoor: Multi-Target Backdoor Attacks on Object Detection
by: Lu, Jialin, et al.
Published: (2025)
by: Lu, Jialin, et al.
Published: (2025)
Naïve Exposure of Generative AI Capabilities Undermines Deepfake Detection
by: Kim, Sunpill, et al.
Published: (2026)
by: Kim, Sunpill, et al.
Published: (2026)
LAID: Lightweight AI-Generated Image Detection in Spatial and Spectral Domains
by: Chivaran, Nicholas, et al.
Published: (2025)
by: Chivaran, Nicholas, et al.
Published: (2025)
Training-Free In-Context Forensic Chain for Image Manipulation Detection and Localization
by: Chen, Rui, et al.
Published: (2025)
by: Chen, Rui, et al.
Published: (2025)
Can VLMs Detect and Localize Fine-Grained AI-Edited Images?
by: Sun, Zhen, et al.
Published: (2025)
by: Sun, Zhen, et al.
Published: (2025)
Similar Items
-
Robust image classification with multi-modal large language models
by: Villani, Francesco, et al.
Published: (2024) -
Machine Learning Security against Data Poisoning: Are We There Yet?
by: Cinà, Antonio Emanuele, et al.
Published: (2022) -
Evaluating the Evaluators: Trust in Adversarial Robustness Tests
by: Cinà, Antonio Emanuele, et al.
Published: (2025) -
Energy-Latency Attacks via Sponge Poisoning
by: Cinà, Antonio Emanuele, et al.
Published: (2022) -
Sonic: Fast and Transferable Data Poisoning on Clustering Algorithms
by: Villani, Francesco, et al.
Published: (2024)