BodhiPromptShield: Pre-Inference Prompt Mediation for Suppressing Privacy Propagation in LLM/VLM Agents
Fuente:
arXiv
Salvato in:
| Autori principali: | Ma, Bo, Wu, Jinsong, Yan, Weiqi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Bodhi VLM: Privacy-Alignment Modeling for Hierarchical Visual Representations in Vision Backbones and VLM Encoders via Bottom-Up and Top-Down Feature Search
di: Ma, Bo, et al.
Pubblicazione: (2026)
di: Ma, Bo, et al.
Pubblicazione: (2026)
REAEDP: Entropy-Calibrated Differentially Private Data Release with Formal Guarantees and Attack-Based Evaluation
di: Ma, Bo, et al.
Pubblicazione: (2026)
di: Ma, Bo, et al.
Pubblicazione: (2026)
UnlearnShield: Shielding Forgotten Privacy against Unlearning Inversion
di: Xue, Lulu, et al.
Pubblicazione: (2026)
di: Xue, Lulu, et al.
Pubblicazione: (2026)
Are You Copying My Prompt? Protecting the Copyright of Vision Prompt for VPaaS via Watermark
di: Ren, Huali, et al.
Pubblicazione: (2024)
di: Ren, Huali, et al.
Pubblicazione: (2024)
Who Can See Through You? Adversarial Shielding Against VLM-Based Attribute Inference Attacks
di: Fan, Yucheng, et al.
Pubblicazione: (2025)
di: Fan, Yucheng, et al.
Pubblicazione: (2025)
PromptSmooth: Certifying Robustness of Medical Vision-Language Models via Prompt Learning
di: Hussein, Noor, et al.
Pubblicazione: (2024)
di: Hussein, Noor, et al.
Pubblicazione: (2024)
AutoMIA: Improved Baselines for Membership Inference Attack via Agentic Self-Exploration
di: Liu, Ruhao, et al.
Pubblicazione: (2026)
di: Liu, Ruhao, et al.
Pubblicazione: (2026)
A Cross-Modal Prompt Injection Attack against Large Vision-Language Models with Image-Only Perturbation
di: Yang, Hao, et al.
Pubblicazione: (2026)
di: Yang, Hao, et al.
Pubblicazione: (2026)
EO-VLM: VLM-Guided Energy Overload Attacks on Vision Models
di: Seo, Minjae, et al.
Pubblicazione: (2025)
di: Seo, Minjae, et al.
Pubblicazione: (2025)
SecureT2I: No More Unauthorized Manipulation on AI Generated Images from Prompts
di: Wu, Xiaodong, et al.
Pubblicazione: (2025)
di: Wu, Xiaodong, et al.
Pubblicazione: (2025)
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
di: Yuan, Lingzhi, et al.
Pubblicazione: (2025)
di: Yuan, Lingzhi, et al.
Pubblicazione: (2025)
Backdoor Attacks on Prompt-Driven Video Segmentation Foundation Models
di: Zhang, Zongmin, et al.
Pubblicazione: (2025)
di: Zhang, Zongmin, et al.
Pubblicazione: (2025)
Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt
di: Ying, Zonghao, et al.
Pubblicazione: (2024)
di: Ying, Zonghao, et al.
Pubblicazione: (2024)
BadBone: Backdoor Attacks Against Backbone Models in Visual Prompt Learning
di: Yang, Ziqing, et al.
Pubblicazione: (2026)
di: Yang, Ziqing, et al.
Pubblicazione: (2026)
Beyond Text Prompts: Precise Concept Erasure through Text-Image Collaboration
di: Li, Jun, et al.
Pubblicazione: (2026)
di: Li, Jun, et al.
Pubblicazione: (2026)
SPARK: Jailbreaking T2V Models by Synergistically Prompting Auditory and Recontextualized Knowledge
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
SurrogatePrompt: Bypassing the Safety Filter of Text-to-Image Models via Substitution
di: Ba, Zhongjie, et al.
Pubblicazione: (2023)
di: Ba, Zhongjie, et al.
Pubblicazione: (2023)
Not All Prompts Are Secure: A Switchable Backdoor Attack Against Pre-trained Vision Transformers
di: Yang, Sheng, et al.
Pubblicazione: (2024)
di: Yang, Sheng, et al.
Pubblicazione: (2024)
Exposing Functional Fusion: A New Class of Strategic Backdoor in Dynamic Prompt Architectures
di: Liu, Zeyao, et al.
Pubblicazione: (2026)
di: Liu, Zeyao, et al.
Pubblicazione: (2026)
Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM Security
di: Fang, Xiang, et al.
Pubblicazione: (2026)
di: Fang, Xiang, et al.
Pubblicazione: (2026)
Improving Adversarial Transferability on Vision Transformers via Forward Propagation Refinement
di: Ren, Yuchen, et al.
Pubblicazione: (2025)
di: Ren, Yuchen, et al.
Pubblicazione: (2025)
ComPrivDet: Efficient Privacy Object Detection in Compressed Domains Through Inference Reuse
di: Yao, Yunhao, et al.
Pubblicazione: (2026)
di: Yao, Yunhao, et al.
Pubblicazione: (2026)
Volley Revolver: A Novel Matrix-Encoding Method for Privacy-Preserving Neural Networks (Inference)
di: Chiang, John
Pubblicazione: (2022)
di: Chiang, John
Pubblicazione: (2022)
Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone Agents
di: Lin, Zhixin, et al.
Pubblicazione: (2025)
di: Lin, Zhixin, et al.
Pubblicazione: (2025)
DAVSP: Safety Alignment for Large Vision-Language Models via Deep Aligned Visual Safety Prompt
di: Zhang, Yitong, et al.
Pubblicazione: (2025)
di: Zhang, Yitong, et al.
Pubblicazione: (2025)
Revisiting the Privacy Risks of Split Inference: A GAN-Based Data Reconstruction Attack via Progressive Feature Optimization
di: Qiu, Yixiang, et al.
Pubblicazione: (2025)
di: Qiu, Yixiang, et al.
Pubblicazione: (2025)
Image-based Prompt Injection: Hijacking Multimodal LLMs through Visually Embedded Adversarial Instructions
di: Nagaraja, Neha, et al.
Pubblicazione: (2026)
di: Nagaraja, Neha, et al.
Pubblicazione: (2026)
One Prompt to Verify Your Models: Black-Box Text-to-Image Models Verification via Non-Transferable Adversarial Attacks
di: Guo, Ji, et al.
Pubblicazione: (2024)
di: Guo, Ji, et al.
Pubblicazione: (2024)
SlowBA: An efficiency backdoor attack towards VLM-based GUI agents
di: Li, Junxian, et al.
Pubblicazione: (2026)
di: Li, Junxian, et al.
Pubblicazione: (2026)
EditShield: Protecting Unauthorized Image Editing by Instruction-guided Diffusion Models
di: Chen, Ruoxi, et al.
Pubblicazione: (2023)
di: Chen, Ruoxi, et al.
Pubblicazione: (2023)
Iteratively Prompting Multimodal LLMs to Reproduce Natural and AI-Generated Images
di: Naseh, Ali, et al.
Pubblicazione: (2024)
di: Naseh, Ali, et al.
Pubblicazione: (2024)
Few-Shot Adversarial Prompt Learning on Vision-Language Models
di: Zhou, Yiwei, et al.
Pubblicazione: (2024)
di: Zhou, Yiwei, et al.
Pubblicazione: (2024)
FT-Shield: A Watermark Against Unauthorized Fine-tuning in Text-to-Image Diffusion Models
di: Cui, Yingqian, et al.
Pubblicazione: (2023)
di: Cui, Yingqian, et al.
Pubblicazione: (2023)
BSPA: Exploring Black-box Stealthy Prompt Attacks against Image Generators
di: Tian, Yu, et al.
Pubblicazione: (2024)
di: Tian, Yu, et al.
Pubblicazione: (2024)
Harnessing Hyperbolic Geometry for Harmful Prompt Detection and Sanitization
di: Maljkovic, Igor, et al.
Pubblicazione: (2026)
di: Maljkovic, Igor, et al.
Pubblicazione: (2026)
GEO-Detective: Unveiling Location Privacy Risks in Images with LLM Agents
di: Zhang, Xinyu, et al.
Pubblicazione: (2025)
di: Zhang, Xinyu, et al.
Pubblicazione: (2025)
Be Careful What You Smooth For: Label Smoothing Can Be a Privacy Shield but Also a Catalyst for Model Inversion Attacks
di: Struppek, Lukas, et al.
Pubblicazione: (2023)
di: Struppek, Lukas, et al.
Pubblicazione: (2023)
VLATTACK: Multimodal Adversarial Attacks on Vision-Language Tasks via Pre-trained Models
di: Yin, Ziyi, et al.
Pubblicazione: (2023)
di: Yin, Ziyi, et al.
Pubblicazione: (2023)
PLA: Prompt Learning Attack against Text-to-Image Generative Models
di: Lyu, Xinqi, et al.
Pubblicazione: (2025)
di: Lyu, Xinqi, et al.
Pubblicazione: (2025)
Fooling the Watchers: Breaking AIGC Detectors via Semantic Prompt Attacks
di: Hao, Run, et al.
Pubblicazione: (2025)
di: Hao, Run, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Bodhi VLM: Privacy-Alignment Modeling for Hierarchical Visual Representations in Vision Backbones and VLM Encoders via Bottom-Up and Top-Down Feature Search
di: Ma, Bo, et al.
Pubblicazione: (2026) -
REAEDP: Entropy-Calibrated Differentially Private Data Release with Formal Guarantees and Attack-Based Evaluation
di: Ma, Bo, et al.
Pubblicazione: (2026) -
UnlearnShield: Shielding Forgotten Privacy against Unlearning Inversion
di: Xue, Lulu, et al.
Pubblicazione: (2026) -
Are You Copying My Prompt? Protecting the Copyright of Vision Prompt for VPaaS via Watermark
di: Ren, Huali, et al.
Pubblicazione: (2024) -
Who Can See Through You? Adversarial Shielding Against VLM-Based Attribute Inference Attacks
di: Fan, Yucheng, et al.
Pubblicazione: (2025)