PAPILLON: Efficient and Stealthy Fuzz Testing-Powered Jailbreaks for LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gong, Xueluan, Li, Mingzhe, Zhang, Yilin, Ran, Fengyuan, Chen, Chen, Chen, Yanjiao, Wang, Qian, Lam, Kwok-Yan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluating and Mitigating LLM-as-a-judge Bias in Communication Systems
von: Gao, Jiaxin, et al.
Veröffentlicht: (2025)
von: Gao, Jiaxin, et al.
Veröffentlicht: (2025)
Beyond Max Tokens: Stealthy Resource Amplification via Tool Calling Chains in LLM Agents
von: Zhou, Kaiyu, et al.
Veröffentlicht: (2026)
von: Zhou, Kaiyu, et al.
Veröffentlicht: (2026)
Not All Tokens Are Created Equal: Query-Efficient Jailbreak Fuzzing for LLMs
von: Chen, Wenyu, et al.
Veröffentlicht: (2026)
von: Chen, Wenyu, et al.
Veröffentlicht: (2026)
ARMOR: Shielding Unlearnable Examples against Data Augmentation
von: Gong, Xueluan, et al.
Veröffentlicht: (2025)
von: Gong, Xueluan, et al.
Veröffentlicht: (2025)
Plato's Form: Toward Backdoor Defense-as-a-Service for LLMs with Prototype Representations
von: Chen, Chen, et al.
Veröffentlicht: (2026)
von: Chen, Chen, et al.
Veröffentlicht: (2026)
TEMPLATEFUZZ: Fine-Grained Chat Template Fuzzing for Jailbreaking and Red Teaming LLMs
von: Shen, Qingchao, et al.
Veröffentlicht: (2026)
von: Shen, Qingchao, et al.
Veröffentlicht: (2026)
Proactive Detection of Physical Inter-rule Vulnerabilities in IoT Services Using a Deep Learning Approach
von: Huang, Bing, et al.
Veröffentlicht: (2024)
von: Huang, Bing, et al.
Veröffentlicht: (2024)
Hiding in Plain Sight: A Steganographic Approach to Stealthy LLM Jailbreaks
von: Geng, Jianing, et al.
Veröffentlicht: (2025)
von: Geng, Jianing, et al.
Veröffentlicht: (2025)
Threats, Attacks, and Defenses in Machine Unlearning: A Survey
von: Liu, Ziyao, et al.
Veröffentlicht: (2024)
von: Liu, Ziyao, et al.
Veröffentlicht: (2024)
PROMPTFUZZ: Harnessing Fuzzing Techniques for Robust Testing of Prompt Injection in LLMs
von: Yu, Jiahao, et al.
Veröffentlicht: (2024)
von: Yu, Jiahao, et al.
Veröffentlicht: (2024)
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations
von: Li, Xiaohu, et al.
Veröffentlicht: (2025)
von: Li, Xiaohu, et al.
Veröffentlicht: (2025)
Fuzz-Testing Meets LLM-Based Agents: An Automated and Efficient Framework for Jailbreaking Text-To-Image Generation Models
von: Dong, Yingkai, et al.
Veröffentlicht: (2024)
von: Dong, Yingkai, et al.
Veröffentlicht: (2024)
Invisible to Humans, Triggered by Agents: Stealthy Jailbreak Attacks on Mobile Vision-Language Agents
von: Ding, Renhua, et al.
Veröffentlicht: (2025)
von: Ding, Renhua, et al.
Veröffentlicht: (2025)
Efficient Privacy-Preserving Retrieval Augmented Generation with Distance-Preserving Encryption
von: Ye, Huanyi, et al.
Veröffentlicht: (2026)
von: Ye, Huanyi, et al.
Veröffentlicht: (2026)
Transferable & Stealthy Ensemble Attacks: A Black-Box Jailbreaking Framework for Large Language Models
von: Yang, Yiqi, et al.
Veröffentlicht: (2024)
von: Yang, Yiqi, et al.
Veröffentlicht: (2024)
Injecting Universal Jailbreak Backdoors into LLMs in Minutes
von: Chen, Zhuowei, et al.
Veröffentlicht: (2025)
von: Chen, Zhuowei, et al.
Veröffentlicht: (2025)
OrchJail: Jailbreaking Tool-Calling Text-to-Image Agents by Orchestration-Guided Fuzzing
von: Chen, Jianming, et al.
Veröffentlicht: (2026)
von: Chen, Jianming, et al.
Veröffentlicht: (2026)
Trojan's Whisper: Stealthy Manipulation of OpenClaw through Injected Bootstrapped Guidance
von: Liu, Fazhong, et al.
Veröffentlicht: (2026)
von: Liu, Fazhong, et al.
Veröffentlicht: (2026)
LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge
von: Li, Songze, et al.
Veröffentlicht: (2025)
von: Li, Songze, et al.
Veröffentlicht: (2025)
SeqAR: Jailbreak LLMs with Sequential Auto-Generated Characters
von: Yang, Yan, et al.
Veröffentlicht: (2024)
von: Yang, Yan, et al.
Veröffentlicht: (2024)
DECEIVE-AFC: Adversarial Claim Attacks against Search-Enabled LLM-based Fact-Checking Systems
von: Ou, Haoran, et al.
Veröffentlicht: (2026)
von: Ou, Haoran, et al.
Veröffentlicht: (2026)
Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction
von: Wang, Hongtao, et al.
Veröffentlicht: (2026)
von: Wang, Hongtao, et al.
Veröffentlicht: (2026)
FlipAttack: Jailbreak LLMs via Flipping
von: Liu, Yue, et al.
Veröffentlicht: (2024)
von: Liu, Yue, et al.
Veröffentlicht: (2024)
Testing Storage-System Correctness: Challenges, Fuzzing Limitations, and AI-Augmented Opportunities
von: Wang, Ying, et al.
Veröffentlicht: (2026)
von: Wang, Ying, et al.
Veröffentlicht: (2026)
Enhancing Jailbreak Attacks on LLMs via Persona Prompts
von: Zhang, Zheng, et al.
Veröffentlicht: (2025)
von: Zhang, Zheng, et al.
Veröffentlicht: (2025)
RAJ-PGA: Reasoning-Activated Jailbreak and Principle-Guided Alignment Framework for Large Reasoning Models
von: Chen, Jianhao, et al.
Veröffentlicht: (2025)
von: Chen, Jianhao, et al.
Veröffentlicht: (2025)
Why Do Aligned LLMs Remain Jailbreakable: Refusal-Escape Directions, Operator-Level Sources, and Safety-Utility Trade-off
von: Chen, Yu, et al.
Veröffentlicht: (2026)
von: Chen, Yu, et al.
Veröffentlicht: (2026)
The Trojan Example: Jailbreaking LLMs through Template Filling and Unsafety Reasoning
von: Liu, Mingrui, et al.
Veröffentlicht: (2025)
von: Liu, Mingrui, et al.
Veröffentlicht: (2025)
CoTSRF: Utilize Chain of Thought as Stealthy and Robust Fingerprint of Large Language Models
von: Ren, Zhenzhen, et al.
Veröffentlicht: (2025)
von: Ren, Zhenzhen, et al.
Veröffentlicht: (2025)
Distract Large Language Models for Automatic Jailbreak Attack
von: Xiao, Zeguan, et al.
Veröffentlicht: (2024)
von: Xiao, Zeguan, et al.
Veröffentlicht: (2024)
Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward
von: Guo, Weiyang, et al.
Veröffentlicht: (2026)
von: Guo, Weiyang, et al.
Veröffentlicht: (2026)
Confusion is the Final Barrier: Rethinking Jailbreak Evaluation and Investigating the Real Misuse Threat of LLMs
von: Yan, Yu, et al.
Veröffentlicht: (2025)
von: Yan, Yu, et al.
Veröffentlicht: (2025)
No Attack Required: Semantic Fuzzing for Specification Violations in Agent Skills
von: Li, Ying, et al.
Veröffentlicht: (2026)
von: Li, Ying, et al.
Veröffentlicht: (2026)
FlowMur: A Stealthy and Practical Audio Backdoor Attack with Limited Knowledge
von: Lan, Jiahe, et al.
Veröffentlicht: (2023)
von: Lan, Jiahe, et al.
Veröffentlicht: (2023)
Graph of Attacks with Pruning: Optimizing Stealthy Jailbreak Prompt Generation for Enhanced LLM Content Moderation
von: Schwartz, Daniel, et al.
Veröffentlicht: (2025)
von: Schwartz, Daniel, et al.
Veröffentlicht: (2025)
JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models
von: Ran, Delong, et al.
Veröffentlicht: (2024)
von: Ran, Delong, et al.
Veröffentlicht: (2024)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
von: Jaiswal, Piyush, et al.
Veröffentlicht: (2026)
von: Jaiswal, Piyush, et al.
Veröffentlicht: (2026)
Re-Triggering Safeguards within LLMs for Jailbreak Detection
von: Lin, Zheng, et al.
Veröffentlicht: (2026)
von: Lin, Zheng, et al.
Veröffentlicht: (2026)
PUZZLED: Jailbreaking LLMs through Word-Based Puzzles
von: Ahn, Yelim, et al.
Veröffentlicht: (2025)
von: Ahn, Yelim, et al.
Veröffentlicht: (2025)
LLMs Caught in the Crossfire: Malware Requests and Jailbreak Challenges
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Evaluating and Mitigating LLM-as-a-judge Bias in Communication Systems
von: Gao, Jiaxin, et al.
Veröffentlicht: (2025) -
Beyond Max Tokens: Stealthy Resource Amplification via Tool Calling Chains in LLM Agents
von: Zhou, Kaiyu, et al.
Veröffentlicht: (2026) -
Not All Tokens Are Created Equal: Query-Efficient Jailbreak Fuzzing for LLMs
von: Chen, Wenyu, et al.
Veröffentlicht: (2026) -
ARMOR: Shielding Unlearnable Examples against Data Augmentation
von: Gong, Xueluan, et al.
Veröffentlicht: (2025) -
Plato's Form: Toward Backdoor Defense-as-a-Service for LLMs with Prototype Representations
von: Chen, Chen, et al.
Veröffentlicht: (2026)