NEXUS: Network Exploration for eXploiting Unsafe Sequences in Multi-Turn LLM Jailbreaks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Asl, Javad Rafiei, Narula, Sidhant, Ghasemigol, Mohammad, Blanco, Eduardo, Takabi, Daniel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models
von: Narula, Sidhant, et al.
Veröffentlicht: (2025)
von: Narula, Sidhant, et al.
Veröffentlicht: (2025)
MCP-DPT: A Defense-Placement Taxonomy and Coverage Analysis for Model Context Protocol Security
von: Rostamzadeh, Mehrdad, et al.
Veröffentlicht: (2026)
von: Rostamzadeh, Mehrdad, et al.
Veröffentlicht: (2026)
AXE: An Agentic eXploit Engine for Confirming Zero-Day Vulnerability Reports
von: Sajadi, Amirali, et al.
Veröffentlicht: (2026)
von: Sajadi, Amirali, et al.
Veröffentlicht: (2026)
SSCAE -- Semantic, Syntactic, and Context-aware natural language Adversarial Examples generator
von: Asl, Javad Rafiei, et al.
Veröffentlicht: (2024)
von: Asl, Javad Rafiei, et al.
Veröffentlicht: (2024)
BioShield: A Context-Aware Firewall for Securing Bio-LLMs
von: Das, Protiva, et al.
Veröffentlicht: (2026)
von: Das, Protiva, et al.
Veröffentlicht: (2026)
The Echo Chamber Multi-Turn LLM Jailbreak
von: Alobaid, Ahmad, et al.
Veröffentlicht: (2026)
von: Alobaid, Ahmad, et al.
Veröffentlicht: (2026)
MOFHEI: Model Optimizing Framework for Fast and Efficient Homomorphically Encrypted Neural Network Inference
von: Ghazvinian, Parsa, et al.
Veröffentlicht: (2024)
von: Ghazvinian, Parsa, et al.
Veröffentlicht: (2024)
Active Honeypot Guardrail System: Probing and Confirming Multi-Turn LLM Jailbreaks
von: Wu, ChenYu, et al.
Veröffentlicht: (2025)
von: Wu, ChenYu, et al.
Veröffentlicht: (2025)
Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack
von: Russinovich, Mark, et al.
Veröffentlicht: (2024)
von: Russinovich, Mark, et al.
Veröffentlicht: (2024)
A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks
von: Bullwinkel, Blake, et al.
Veröffentlicht: (2025)
von: Bullwinkel, Blake, et al.
Veröffentlicht: (2025)
Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs
von: Xing, Wenpeng, et al.
Veröffentlicht: (2025)
von: Xing, Wenpeng, et al.
Veröffentlicht: (2025)
ICON: Intent-Context Coupling for Efficient Multi-Turn Jailbreak Attack
von: Lin, Xingwei, et al.
Veröffentlicht: (2026)
von: Lin, Xingwei, et al.
Veröffentlicht: (2026)
MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks
von: Zhang, Xinkai, et al.
Veröffentlicht: (2026)
von: Zhang, Xinkai, et al.
Veröffentlicht: (2026)
Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
von: Tong, Haibo, et al.
Veröffentlicht: (2025)
von: Tong, Haibo, et al.
Veröffentlicht: (2025)
I can't see it but I can Fine-tune it: On Encrypted Fine-tuning of Transformers using Fully Homomorphic Encryption
von: Panzade, Prajwal, et al.
Veröffentlicht: (2024)
von: Panzade, Prajwal, et al.
Veröffentlicht: (2024)
Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks
von: Gibbs, Tom, et al.
Veröffentlicht: (2024)
von: Gibbs, Tom, et al.
Veröffentlicht: (2024)
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
Tempest: Autonomous Multi-Turn Jailbreaking of Large Language Models with Tree Search
von: Zhou, Andy, et al.
Veröffentlicht: (2025)
von: Zhou, Andy, et al.
Veröffentlicht: (2025)
CryptoAnalystBench: Failures in Multi-Tool Long-Form LLM Analysis
von: Eswaran, Anushri, et al.
Veröffentlicht: (2026)
von: Eswaran, Anushri, et al.
Veröffentlicht: (2026)
Amplified Vulnerabilities: Structured Jailbreak Attacks on LLM-based Multi-Agent Debate
von: Qi, Senmao, et al.
Veröffentlicht: (2025)
von: Qi, Senmao, et al.
Veröffentlicht: (2025)
Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models
von: Nihal, Ragib Amin, et al.
Veröffentlicht: (2025)
von: Nihal, Ragib Amin, et al.
Veröffentlicht: (2025)
SoK: Privacy Preserving Machine Learning using Functional Encryption: Opportunities and Challenges
von: Panzade, Prajwal, et al.
Veröffentlicht: (2022)
von: Panzade, Prajwal, et al.
Veröffentlicht: (2022)
The Great Pretender: A Stochasticity Problem in LLM Jailbreak
von: Monteuuis, Jean-Philippe, et al.
Veröffentlicht: (2026)
von: Monteuuis, Jean-Philippe, et al.
Veröffentlicht: (2026)
Heterogeneous Graph Backdoor Attack
von: Chen, Jiawei, et al.
Veröffentlicht: (2025)
von: Chen, Jiawei, et al.
Veröffentlicht: (2025)
Latent Adversarial Detection: Adaptive Probing of LLM Activations for Multi-Turn Attack Detection
von: Kulkarni, Prashant
Veröffentlicht: (2026)
von: Kulkarni, Prashant
Veröffentlicht: (2026)
Behind the Mask: Benchmarking Camouflaged Jailbreaks in Large Language Models
von: Zheng, Youjia, et al.
Veröffentlicht: (2025)
von: Zheng, Youjia, et al.
Veröffentlicht: (2025)
AutoDAN-Reasoning: Enhancing Strategies Exploration based Jailbreak Attacks with Test-Time Scaling
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2025)
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2025)
Formalization Driven LLM Prompt Jailbreaking via Reinforcement Learning
von: Wang, Zhaoqi, et al.
Veröffentlicht: (2025)
von: Wang, Zhaoqi, et al.
Veröffentlicht: (2025)
Siren: A Learning-Based Multi-Turn Attack Framework for Simulating Real-World Human Jailbreak Behaviors
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities
von: Qu, Yiting, et al.
Veröffentlicht: (2025)
von: Qu, Yiting, et al.
Veröffentlicht: (2025)
Improving LLM Outputs Against Jailbreak Attacks with Expert Model Integration
von: Tsmindashvili, Tatia, et al.
Veröffentlicht: (2025)
von: Tsmindashvili, Tatia, et al.
Veröffentlicht: (2025)
Hiding in Plain Sight: A Steganographic Approach to Stealthy LLM Jailbreaks
von: Geng, Jianing, et al.
Veröffentlicht: (2025)
von: Geng, Jianing, et al.
Veröffentlicht: (2025)
Multi-turn Jailbreaking Attack in Multi-Modal Large Language Models
von: Das, Badhan Chandra, et al.
Veröffentlicht: (2026)
von: Das, Badhan Chandra, et al.
Veröffentlicht: (2026)
Transient Turn Injection: Exposing Stateless Multi-Turn Vulnerabilities in Large Language Models
von: Rayhan, Naheed, et al.
Veröffentlicht: (2026)
von: Rayhan, Naheed, et al.
Veröffentlicht: (2026)
X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents
von: Rahman, Salman, et al.
Veröffentlicht: (2025)
von: Rahman, Salman, et al.
Veröffentlicht: (2025)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
von: Reddy, Aashray, et al.
Veröffentlicht: (2025)
von: Reddy, Aashray, et al.
Veröffentlicht: (2025)
Unsafe LLM-Based Search: Quantitative Analysis and Mitigation of Safety Risks in AI Web Search
von: Luo, Zeren, et al.
Veröffentlicht: (2025)
von: Luo, Zeren, et al.
Veröffentlicht: (2025)
Bleeding Pathways: Vanishing Discriminability in LLM Hidden States Fuels Jailbreak Attacks
von: Zhang, Yingjie, et al.
Veröffentlicht: (2025)
von: Zhang, Yingjie, et al.
Veröffentlicht: (2025)
MetaBreak: Jailbreaking Online LLM Services via Special Token Manipulation
von: Zhu, Wentian, et al.
Veröffentlicht: (2025)
von: Zhu, Wentian, et al.
Veröffentlicht: (2025)
The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring
von: Hossain, Ismail, et al.
Veröffentlicht: (2026)
von: Hossain, Ismail, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models
von: Narula, Sidhant, et al.
Veröffentlicht: (2025) -
MCP-DPT: A Defense-Placement Taxonomy and Coverage Analysis for Model Context Protocol Security
von: Rostamzadeh, Mehrdad, et al.
Veröffentlicht: (2026) -
AXE: An Agentic eXploit Engine for Confirming Zero-Day Vulnerability Reports
von: Sajadi, Amirali, et al.
Veröffentlicht: (2026) -
SSCAE -- Semantic, Syntactic, and Context-aware natural language Adversarial Examples generator
von: Asl, Javad Rafiei, et al.
Veröffentlicht: (2024) -
BioShield: A Context-Aware Firewall for Securing Bio-LLMs
von: Das, Protiva, et al.
Veröffentlicht: (2026)