Prompt, Divide, and Conquer: Bypassing Large Language Model Safety Filters via Segmented and Distributed Prompt Processing
Fuente:
arXiv
Saved in:
| Main Authors: | Wahréus, Johan, Hussain, Ahmed, Papadimitratos, Panos |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CySecBench: Generative AI-based CyberSecurity-focused Prompt Dataset for Benchmarking Large Language Models
by: Wahréus, Johan, et al.
Published: (2025)
by: Wahréus, Johan, et al.
Published: (2025)
Jailbreaking Large Language Models Through Content Concretization
by: Wahréus, Johan, et al.
Published: (2025)
by: Wahréus, Johan, et al.
Published: (2025)
LLMs Have Rhythm: Fingerprinting Large Language Models Using Inter-Token Times and Network Traffic Analysis
by: Alhazbi, Saeif, et al.
Published: (2025)
by: Alhazbi, Saeif, et al.
Published: (2025)
Attention in Motion: Secure Platooning via Transformer-based Misbehavior Detection
by: Kalogiannis, Konstantinos, et al.
Published: (2025)
by: Kalogiannis, Konstantinos, et al.
Published: (2025)
Less Data, More Security: Advancing Cybersecurity LLMs Specialization via Resource-Efficient Domain-Adaptive Continuous Pre-training with Minimal Tokens
by: Salahuddin, Salahuddin, et al.
Published: (2025)
by: Salahuddin, Salahuddin, et al.
Published: (2025)
Edge AI-based Radio Frequency Fingerprinting for IoT Networks
by: Hussain, Ahmed Mohamed, et al.
Published: (2024)
by: Hussain, Ahmed Mohamed, et al.
Published: (2024)
FedTrident: Resilient Road Condition Classification Against Poisoning Attacks in Federated Learning
by: Liu, Sheng, et al.
Published: (2026)
by: Liu, Sheng, et al.
Published: (2026)
Prompt Injection Attacks on Large Language Models in Oncology
by: Clusmann, Jan, et al.
Published: (2024)
by: Clusmann, Jan, et al.
Published: (2024)
SequentialBreak: Large Language Models Can be Fooled by Embedding Jailbreak Prompts into Sequential Prompt Chains
by: Saiem, Bijoy Ahmed, et al.
Published: (2024)
by: Saiem, Bijoy Ahmed, et al.
Published: (2024)
PLeak: Prompt Leaking Attacks against Large Language Model Applications
by: Hui, Bo, et al.
Published: (2024)
by: Hui, Bo, et al.
Published: (2024)
DEFEND: Poisoned Model Detection and Malicious Client Exclusion Mechanism for Secure Federated Learning-based Road Condition Classification
by: Liu, Sheng, et al.
Published: (2025)
by: Liu, Sheng, et al.
Published: (2025)
Safeguarding Federated Learning-based Road Condition Classification
by: Liu, Sheng, et al.
Published: (2025)
by: Liu, Sheng, et al.
Published: (2025)
PAMPOS: Causal Transformer-based Trajectory Prediction for Attack-Agnostic Misbehavior Detection in V2X Networks
by: Kalogiannis, Konstantinos, et al.
Published: (2026)
by: Kalogiannis, Konstantinos, et al.
Published: (2026)
LMEraser: Large Model Unlearning through Adaptive Prompt Tuning
by: Xu, Jie, et al.
Published: (2024)
by: Xu, Jie, et al.
Published: (2024)
Attention Eclipse: Manipulating Attention to Bypass LLM Safety-Alignment
by: Zaree, Pedram, et al.
Published: (2025)
by: Zaree, Pedram, et al.
Published: (2025)
Unlocking Memorization in Large Language Models with Dynamic Soft Prompting
by: Wang, Zhepeng, et al.
Published: (2024)
by: Wang, Zhepeng, et al.
Published: (2024)
Radio Frequency Fingerprinting via Deep Learning: Challenges and Opportunities
by: Al-Hazbi, Saeif, et al.
Published: (2023)
by: Al-Hazbi, Saeif, et al.
Published: (2023)
An Empirical Evaluation of LLM-Generated Code Security Across Prompting Methods
by: Kharma, Mohammed, et al.
Published: (2026)
by: Kharma, Mohammed, et al.
Published: (2026)
Understanding the Effects of Safety Unalignment on Large Language Models
by: Halloran, John T.
Published: (2026)
by: Halloran, John T.
Published: (2026)
Certifying LLM Safety against Adversarial Prompting
by: Kumar, Aounon, et al.
Published: (2023)
by: Kumar, Aounon, et al.
Published: (2023)
Arondight: Red Teaming Large Vision Language Models with Auto-generated Multi-modal Jailbreak Prompts
by: Liu, Yi, et al.
Published: (2024)
by: Liu, Yi, et al.
Published: (2024)
RECAP: A Resource-Efficient Method for Adversarial Prompting in Large Language Models
by: Chugh, Rishit
Published: (2026)
by: Chugh, Rishit
Published: (2026)
Using Hallucinations to Bypass GPT4's Filter
by: Lemkin, Benjamin
Published: (2024)
by: Lemkin, Benjamin
Published: (2024)
AttentionGuard: Transformer-based Misbehavior Detection for Secure Vehicular Platoons
by: Li, Hexu, et al.
Published: (2025)
by: Li, Hexu, et al.
Published: (2025)
SWAP: Towards Copyright Auditing of Soft Prompts via Sequential Watermarking
by: Yang, Wenyuan, et al.
Published: (2025)
by: Yang, Wenyuan, et al.
Published: (2025)
No, of Course I Can! Deeper Fine-Tuning Attacks That Bypass Token-Level Safety Mechanisms
by: Kazdan, Joshua, et al.
Published: (2025)
by: Kazdan, Joshua, et al.
Published: (2025)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
by: Reddy, Aashray, et al.
Published: (2025)
by: Reddy, Aashray, et al.
Published: (2025)
How Not to Detect Prompt Injections with an LLM
by: Choudhary, Sarthak, et al.
Published: (2025)
by: Choudhary, Sarthak, et al.
Published: (2025)
Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts
by: Wu, Yuanwei, et al.
Published: (2023)
by: Wu, Yuanwei, et al.
Published: (2023)
Bypassing Prompt Guards in Production with Controlled-Release Prompting
by: Fairoze, Jaiden, et al.
Published: (2025)
by: Fairoze, Jaiden, et al.
Published: (2025)
UpSafe$^\circ$C: Upcycling for Controllable Safety in Large Language Models
by: Sun, Yuhao, et al.
Published: (2025)
by: Sun, Yuhao, et al.
Published: (2025)
Knowing without Acting: The Disentangled Geometry of Safety Mechanisms in Large Language Models
by: Wu, Jinman, et al.
Published: (2026)
by: Wu, Jinman, et al.
Published: (2026)
Harry Potter is Still Here! Probing Knowledge Leakage in Targeted Unlearned Large Language Models via Automated Adversarial Prompting
by: To, Bang Trinh Tran, et al.
Published: (2025)
by: To, Bang Trinh Tran, et al.
Published: (2025)
PISanitizer: Preventing Prompt Injection to Long-Context LLMs via Prompt Sanitization
by: Geng, Runpeng, et al.
Published: (2025)
by: Geng, Runpeng, et al.
Published: (2025)
Logic layer Prompt Control Injection (LPCI): A Novel Security Vulnerability Class in Agentic Systems
by: Atta, Hammad, et al.
Published: (2025)
by: Atta, Hammad, et al.
Published: (2025)
Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation
by: Huang, Tiansheng, et al.
Published: (2025)
by: Huang, Tiansheng, et al.
Published: (2025)
Attention Tracker: Detecting Prompt Injection Attacks in LLMs
by: Hung, Kuo-Han, et al.
Published: (2024)
by: Hung, Kuo-Han, et al.
Published: (2024)
Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning
by: Liu, Guozhi, et al.
Published: (2025)
by: Liu, Guozhi, et al.
Published: (2025)
Refusing Safe Prompts for Multi-modal Large Language Models
by: Shao, Zedian, et al.
Published: (2024)
by: Shao, Zedian, et al.
Published: (2024)
Bypassing LLM Guardrails: An Empirical Analysis of Evasion Attacks against Prompt Injection and Jailbreak Detection Systems
by: Hackett, William, et al.
Published: (2025)
by: Hackett, William, et al.
Published: (2025)
Similar Items
-
CySecBench: Generative AI-based CyberSecurity-focused Prompt Dataset for Benchmarking Large Language Models
by: Wahréus, Johan, et al.
Published: (2025) -
Jailbreaking Large Language Models Through Content Concretization
by: Wahréus, Johan, et al.
Published: (2025) -
LLMs Have Rhythm: Fingerprinting Large Language Models Using Inter-Token Times and Network Traffic Analysis
by: Alhazbi, Saeif, et al.
Published: (2025) -
Attention in Motion: Secure Platooning via Transformer-based Misbehavior Detection
by: Kalogiannis, Konstantinos, et al.
Published: (2025) -
Less Data, More Security: Advancing Cybersecurity LLMs Specialization via Resource-Efficient Domain-Adaptive Continuous Pre-training with Minimal Tokens
by: Salahuddin, Salahuddin, et al.
Published: (2025)