`Do as I say not as I do': A Semi-Automated Approach for Jailbreak Prompt Attack against Multimodal LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chiu, Chun Wai, Huang, Linghan, Li, Bo, Chen, Huaming, Choo, Kim-Kwang Raymond |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Lightweight Vulnerability Detection from Code Metrics and Token Features
von: Chiu, Chun Yin
Veröffentlicht: (2026)
von: Chiu, Chun Yin
Veröffentlicht: (2026)
Automated Attack Synthesis for Constant Product Market Makers
von: Han, Sujin, et al.
Veröffentlicht: (2024)
von: Han, Sujin, et al.
Veröffentlicht: (2024)
Compositional Jailbreaking: An Empirical Analysis of Mutator Chain Interactions in Aligned LLMs
von: Bugnot, Reinelle Jan, et al.
Veröffentlicht: (2026)
von: Bugnot, Reinelle Jan, et al.
Veröffentlicht: (2026)
DecipherGuard: Understanding and Deciphering Jailbreak Prompts for a Safer Deployment of Intelligent Software Systems
von: Yang, Rui, et al.
Veröffentlicht: (2025)
von: Yang, Rui, et al.
Veröffentlicht: (2025)
Demonstration Attack against In-Context Learning for Code Intelligence
von: Ge, Yifei, et al.
Veröffentlicht: (2024)
von: Ge, Yifei, et al.
Veröffentlicht: (2024)
Large Language Models for Code Analysis: Do LLMs Really Do Their Job?
von: Fang, Chongzhou, et al.
Veröffentlicht: (2023)
von: Fang, Chongzhou, et al.
Veröffentlicht: (2023)
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
von: Xiong, Chen, et al.
Veröffentlicht: (2024)
von: Xiong, Chen, et al.
Veröffentlicht: (2024)
A Practical Adversarial Attack against Sequence-based Deep Learning Malware Classifiers
von: Tan, Kai, et al.
Veröffentlicht: (2025)
von: Tan, Kai, et al.
Veröffentlicht: (2025)
Defending Code Language Models against Backdoor Attacks with Deceptive Cross-Entropy Loss
von: Yang, Guang, et al.
Veröffentlicht: (2024)
von: Yang, Guang, et al.
Veröffentlicht: (2024)
PrediQL: Automated Testing of GraphQL APIs with LLMs
von: Liu, Shaolun, et al.
Veröffentlicht: (2025)
von: Liu, Shaolun, et al.
Veröffentlicht: (2025)
"Your AI, My Shell": Demystifying Prompt Injection Attacks on Agentic AI Coding Editors
von: Liu, Yue, et al.
Veröffentlicht: (2025)
von: Liu, Yue, et al.
Veröffentlicht: (2025)
Do Fine-Tuned LLMs Understand Vulnerabilities? An Investigation into the Semantic Trap
von: Huang, Feiyang, et al.
Veröffentlicht: (2026)
von: Huang, Feiyang, et al.
Veröffentlicht: (2026)
Automatic Attack Script Generation: a MDA Approach
von: Goux, Quentin, et al.
Veröffentlicht: (2026)
von: Goux, Quentin, et al.
Veröffentlicht: (2026)
From Transactions to Exploits: Automated PoC Synthesis for Real-World DeFi Attacks
von: Su, Xing, et al.
Veröffentlicht: (2026)
von: Su, Xing, et al.
Veröffentlicht: (2026)
APPATCH: Automated Adaptive Prompting Large Language Models for Real-World Software Vulnerability Patching
von: Nong, Yu, et al.
Veröffentlicht: (2024)
von: Nong, Yu, et al.
Veröffentlicht: (2024)
Pinning Is Futile: You Need More Than Local Dependency Versioning to Defend against Supply Chain Attacks
von: He, Hao, et al.
Veröffentlicht: (2025)
von: He, Hao, et al.
Veröffentlicht: (2025)
Automated TEE Adaptation with LLMs: Identifying, Transforming, and Porting Sensitive Functions in Programs
von: Han, Ruidong, et al.
Veröffentlicht: (2025)
von: Han, Ruidong, et al.
Veröffentlicht: (2025)
TEMPLATEFUZZ: Fine-Grained Chat Template Fuzzing for Jailbreaking and Red Teaming LLMs
von: Shen, Qingchao, et al.
Veröffentlicht: (2026)
von: Shen, Qingchao, et al.
Veröffentlicht: (2026)
Can I Check What I Designed? Mapping Security Design DSLs to Code Analyzers
von: Peldszus, Sven, et al.
Veröffentlicht: (2026)
von: Peldszus, Sven, et al.
Veröffentlicht: (2026)
Evaluating and Improving the Robustness of Security Attack Detectors Generated by LLMs
von: Pasini, Samuele, et al.
Veröffentlicht: (2024)
von: Pasini, Samuele, et al.
Veröffentlicht: (2024)
PolyJailbreak: Cross-Modal Jailbreaking Attacks on Black-Box Multimodal LLMs
von: Wang, Xinkai, et al.
Veröffentlicht: (2025)
von: Wang, Xinkai, et al.
Veröffentlicht: (2025)
XOXO: Stealthy Cross-Origin Context Poisoning Attacks against AI Coding Assistants
von: Štorek, Adam, et al.
Veröffentlicht: (2025)
von: Štorek, Adam, et al.
Veröffentlicht: (2025)
{A New Hope}: Contextual Privacy Policies for Mobile Applications and An Approach Toward Automated Generation
von: Pan, Shidong, et al.
Veröffentlicht: (2024)
von: Pan, Shidong, et al.
Veröffentlicht: (2024)
Unknown Attack Detection in IoT Networks using Large Language Models: A Robust, Data-efficient Approach
von: Ali, Shan, et al.
Veröffentlicht: (2026)
von: Ali, Shan, et al.
Veröffentlicht: (2026)
A Taxonomy of System-Level Attacks on Deep Learning Models in Autonomous Vehicles
von: Tehrani, Masoud Jamshidiyan, et al.
Veröffentlicht: (2024)
von: Tehrani, Masoud Jamshidiyan, et al.
Veröffentlicht: (2024)
An LLM-Assisted Easy-to-Trigger Backdoor Attack on Code Completion Models: Injecting Disguised Vulnerabilities against Strong Detection
von: Yan, Shenao, et al.
Veröffentlicht: (2024)
von: Yan, Shenao, et al.
Veröffentlicht: (2024)
AUTOVR: Automated UI Exploration for Detecting Sensitive Data Flow Exposures in Virtual Reality Apps
von: Kim, John Y., et al.
Veröffentlicht: (2025)
von: Kim, John Y., et al.
Veröffentlicht: (2025)
Give LLMs a Security Course: Securing Retrieval-Augmented Code Generation via Knowledge Injection
von: Lin, Bo, et al.
Veröffentlicht: (2025)
von: Lin, Bo, et al.
Veröffentlicht: (2025)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
von: Jaiswal, Piyush, et al.
Veröffentlicht: (2026)
von: Jaiswal, Piyush, et al.
Veröffentlicht: (2026)
Enhancing Jailbreak Attacks on LLMs via Persona Prompts
von: Zhang, Zheng, et al.
Veröffentlicht: (2025)
von: Zhang, Zheng, et al.
Veröffentlicht: (2025)
I Know Who Clones Your Code: Interpretable Smart Contract Similarity Detection
von: Liu, Zhenguang, et al.
Veröffentlicht: (2025)
von: Liu, Zhenguang, et al.
Veröffentlicht: (2025)
Minimal Prompt Perturbations Lead to Code Vulnerabilities: Prompt Fragility and Hidden-State Signals in Coding LLMs
von: Sternfeld, Alexander, et al.
Veröffentlicht: (2026)
von: Sternfeld, Alexander, et al.
Veröffentlicht: (2026)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
von: Ma, Jiachen, et al.
Veröffentlicht: (2024)
von: Ma, Jiachen, et al.
Veröffentlicht: (2024)
Jailbreak Distillation: Renewable Safety Benchmarking
von: Zhang, Jingyu, et al.
Veröffentlicht: (2025)
von: Zhang, Jingyu, et al.
Veröffentlicht: (2025)
Integrating APK Image and Text Data for Enhanced Threat Detection: A Multimodal Deep Learning Approach to Android Malware
von: Arifin, Md Mashrur, et al.
Veröffentlicht: (2026)
von: Arifin, Md Mashrur, et al.
Veröffentlicht: (2026)
Prompt Fuzzing for Fuzz Driver Generation
von: Lyu, Yunlong, et al.
Veröffentlicht: (2023)
von: Lyu, Yunlong, et al.
Veröffentlicht: (2023)
Does Teaming-Up LLMs Improve Secure Code Generation? A Comprehensive Evaluation with Multi-LLMSecCodeEval
von: Sabir, Bushra, et al.
Veröffentlicht: (2026)
von: Sabir, Bushra, et al.
Veröffentlicht: (2026)
"I Don't Use AI for Everything": Exploring Utility, Attitude, and Responsibility of AI-empowered Tools in Software Development
von: Pan, Shidong, et al.
Veröffentlicht: (2024)
von: Pan, Shidong, et al.
Veröffentlicht: (2024)
RL-JACK: Reinforcement Learning-powered Black-box Jailbreaking Attack against LLMs
von: Chen, Xuan, et al.
Veröffentlicht: (2024)
von: Chen, Xuan, et al.
Veröffentlicht: (2024)
Automated Generation of Cybersecurity Exercise Scenarios
von: Skandylas, Charilaos, et al.
Veröffentlicht: (2026)
von: Skandylas, Charilaos, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Lightweight Vulnerability Detection from Code Metrics and Token Features
von: Chiu, Chun Yin
Veröffentlicht: (2026) -
Automated Attack Synthesis for Constant Product Market Makers
von: Han, Sujin, et al.
Veröffentlicht: (2024) -
Compositional Jailbreaking: An Empirical Analysis of Mutator Chain Interactions in Aligned LLMs
von: Bugnot, Reinelle Jan, et al.
Veröffentlicht: (2026) -
DecipherGuard: Understanding and Deciphering Jailbreak Prompts for a Safer Deployment of Intelligent Software Systems
von: Yang, Rui, et al.
Veröffentlicht: (2025) -
Demonstration Attack against In-Context Learning for Code Intelligence
von: Ge, Yifei, et al.
Veröffentlicht: (2024)