BodhiPromptShield: Pre-Inference Prompt Mediation for Suppressing Privacy Propagation in LLM/VLM Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ma, Bo, Wu, Jinsong, Yan, Weiqi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Bodhi VLM: Privacy-Alignment Modeling for Hierarchical Visual Representations in Vision Backbones and VLM Encoders via Bottom-Up and Top-Down Feature Search
von: Ma, Bo, et al.
Veröffentlicht: (2026)
von: Ma, Bo, et al.
Veröffentlicht: (2026)
REAEDP: Entropy-Calibrated Differentially Private Data Release with Formal Guarantees and Attack-Based Evaluation
von: Ma, Bo, et al.
Veröffentlicht: (2026)
von: Ma, Bo, et al.
Veröffentlicht: (2026)
UnlearnShield: Shielding Forgotten Privacy against Unlearning Inversion
von: Xue, Lulu, et al.
Veröffentlicht: (2026)
von: Xue, Lulu, et al.
Veröffentlicht: (2026)
Are You Copying My Prompt? Protecting the Copyright of Vision Prompt for VPaaS via Watermark
von: Ren, Huali, et al.
Veröffentlicht: (2024)
von: Ren, Huali, et al.
Veröffentlicht: (2024)
Who Can See Through You? Adversarial Shielding Against VLM-Based Attribute Inference Attacks
von: Fan, Yucheng, et al.
Veröffentlicht: (2025)
von: Fan, Yucheng, et al.
Veröffentlicht: (2025)
PromptSmooth: Certifying Robustness of Medical Vision-Language Models via Prompt Learning
von: Hussein, Noor, et al.
Veröffentlicht: (2024)
von: Hussein, Noor, et al.
Veröffentlicht: (2024)
AutoMIA: Improved Baselines for Membership Inference Attack via Agentic Self-Exploration
von: Liu, Ruhao, et al.
Veröffentlicht: (2026)
von: Liu, Ruhao, et al.
Veröffentlicht: (2026)
A Cross-Modal Prompt Injection Attack against Large Vision-Language Models with Image-Only Perturbation
von: Yang, Hao, et al.
Veröffentlicht: (2026)
von: Yang, Hao, et al.
Veröffentlicht: (2026)
EO-VLM: VLM-Guided Energy Overload Attacks on Vision Models
von: Seo, Minjae, et al.
Veröffentlicht: (2025)
von: Seo, Minjae, et al.
Veröffentlicht: (2025)
SecureT2I: No More Unauthorized Manipulation on AI Generated Images from Prompts
von: Wu, Xiaodong, et al.
Veröffentlicht: (2025)
von: Wu, Xiaodong, et al.
Veröffentlicht: (2025)
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
von: Yuan, Lingzhi, et al.
Veröffentlicht: (2025)
von: Yuan, Lingzhi, et al.
Veröffentlicht: (2025)
Backdoor Attacks on Prompt-Driven Video Segmentation Foundation Models
von: Zhang, Zongmin, et al.
Veröffentlicht: (2025)
von: Zhang, Zongmin, et al.
Veröffentlicht: (2025)
Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
BadBone: Backdoor Attacks Against Backbone Models in Visual Prompt Learning
von: Yang, Ziqing, et al.
Veröffentlicht: (2026)
von: Yang, Ziqing, et al.
Veröffentlicht: (2026)
Beyond Text Prompts: Precise Concept Erasure through Text-Image Collaboration
von: Li, Jun, et al.
Veröffentlicht: (2026)
von: Li, Jun, et al.
Veröffentlicht: (2026)
SPARK: Jailbreaking T2V Models by Synergistically Prompting Auditory and Recontextualized Knowledge
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
SurrogatePrompt: Bypassing the Safety Filter of Text-to-Image Models via Substitution
von: Ba, Zhongjie, et al.
Veröffentlicht: (2023)
von: Ba, Zhongjie, et al.
Veröffentlicht: (2023)
Not All Prompts Are Secure: A Switchable Backdoor Attack Against Pre-trained Vision Transformers
von: Yang, Sheng, et al.
Veröffentlicht: (2024)
von: Yang, Sheng, et al.
Veröffentlicht: (2024)
Exposing Functional Fusion: A New Class of Strategic Backdoor in Dynamic Prompt Architectures
von: Liu, Zeyao, et al.
Veröffentlicht: (2026)
von: Liu, Zeyao, et al.
Veröffentlicht: (2026)
Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM Security
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
Improving Adversarial Transferability on Vision Transformers via Forward Propagation Refinement
von: Ren, Yuchen, et al.
Veröffentlicht: (2025)
von: Ren, Yuchen, et al.
Veröffentlicht: (2025)
ComPrivDet: Efficient Privacy Object Detection in Compressed Domains Through Inference Reuse
von: Yao, Yunhao, et al.
Veröffentlicht: (2026)
von: Yao, Yunhao, et al.
Veröffentlicht: (2026)
Volley Revolver: A Novel Matrix-Encoding Method for Privacy-Preserving Neural Networks (Inference)
von: Chiang, John
Veröffentlicht: (2022)
von: Chiang, John
Veröffentlicht: (2022)
Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone Agents
von: Lin, Zhixin, et al.
Veröffentlicht: (2025)
von: Lin, Zhixin, et al.
Veröffentlicht: (2025)
DAVSP: Safety Alignment for Large Vision-Language Models via Deep Aligned Visual Safety Prompt
von: Zhang, Yitong, et al.
Veröffentlicht: (2025)
von: Zhang, Yitong, et al.
Veröffentlicht: (2025)
Revisiting the Privacy Risks of Split Inference: A GAN-Based Data Reconstruction Attack via Progressive Feature Optimization
von: Qiu, Yixiang, et al.
Veröffentlicht: (2025)
von: Qiu, Yixiang, et al.
Veröffentlicht: (2025)
Image-based Prompt Injection: Hijacking Multimodal LLMs through Visually Embedded Adversarial Instructions
von: Nagaraja, Neha, et al.
Veröffentlicht: (2026)
von: Nagaraja, Neha, et al.
Veröffentlicht: (2026)
One Prompt to Verify Your Models: Black-Box Text-to-Image Models Verification via Non-Transferable Adversarial Attacks
von: Guo, Ji, et al.
Veröffentlicht: (2024)
von: Guo, Ji, et al.
Veröffentlicht: (2024)
SlowBA: An efficiency backdoor attack towards VLM-based GUI agents
von: Li, Junxian, et al.
Veröffentlicht: (2026)
von: Li, Junxian, et al.
Veröffentlicht: (2026)
EditShield: Protecting Unauthorized Image Editing by Instruction-guided Diffusion Models
von: Chen, Ruoxi, et al.
Veröffentlicht: (2023)
von: Chen, Ruoxi, et al.
Veröffentlicht: (2023)
Iteratively Prompting Multimodal LLMs to Reproduce Natural and AI-Generated Images
von: Naseh, Ali, et al.
Veröffentlicht: (2024)
von: Naseh, Ali, et al.
Veröffentlicht: (2024)
Few-Shot Adversarial Prompt Learning on Vision-Language Models
von: Zhou, Yiwei, et al.
Veröffentlicht: (2024)
von: Zhou, Yiwei, et al.
Veröffentlicht: (2024)
FT-Shield: A Watermark Against Unauthorized Fine-tuning in Text-to-Image Diffusion Models
von: Cui, Yingqian, et al.
Veröffentlicht: (2023)
von: Cui, Yingqian, et al.
Veröffentlicht: (2023)
BSPA: Exploring Black-box Stealthy Prompt Attacks against Image Generators
von: Tian, Yu, et al.
Veröffentlicht: (2024)
von: Tian, Yu, et al.
Veröffentlicht: (2024)
Harnessing Hyperbolic Geometry for Harmful Prompt Detection and Sanitization
von: Maljkovic, Igor, et al.
Veröffentlicht: (2026)
von: Maljkovic, Igor, et al.
Veröffentlicht: (2026)
GEO-Detective: Unveiling Location Privacy Risks in Images with LLM Agents
von: Zhang, Xinyu, et al.
Veröffentlicht: (2025)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2025)
Be Careful What You Smooth For: Label Smoothing Can Be a Privacy Shield but Also a Catalyst for Model Inversion Attacks
von: Struppek, Lukas, et al.
Veröffentlicht: (2023)
von: Struppek, Lukas, et al.
Veröffentlicht: (2023)
VLATTACK: Multimodal Adversarial Attacks on Vision-Language Tasks via Pre-trained Models
von: Yin, Ziyi, et al.
Veröffentlicht: (2023)
von: Yin, Ziyi, et al.
Veröffentlicht: (2023)
PLA: Prompt Learning Attack against Text-to-Image Generative Models
von: Lyu, Xinqi, et al.
Veröffentlicht: (2025)
von: Lyu, Xinqi, et al.
Veröffentlicht: (2025)
Fooling the Watchers: Breaking AIGC Detectors via Semantic Prompt Attacks
von: Hao, Run, et al.
Veröffentlicht: (2025)
von: Hao, Run, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Bodhi VLM: Privacy-Alignment Modeling for Hierarchical Visual Representations in Vision Backbones and VLM Encoders via Bottom-Up and Top-Down Feature Search
von: Ma, Bo, et al.
Veröffentlicht: (2026) -
REAEDP: Entropy-Calibrated Differentially Private Data Release with Formal Guarantees and Attack-Based Evaluation
von: Ma, Bo, et al.
Veröffentlicht: (2026) -
UnlearnShield: Shielding Forgotten Privacy against Unlearning Inversion
von: Xue, Lulu, et al.
Veröffentlicht: (2026) -
Are You Copying My Prompt? Protecting the Copyright of Vision Prompt for VPaaS via Watermark
von: Ren, Huali, et al.
Veröffentlicht: (2024) -
Who Can See Through You? Adversarial Shielding Against VLM-Based Attribute Inference Attacks
von: Fan, Yucheng, et al.
Veröffentlicht: (2025)