Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ying, Zonghao, Zhang, Deyue, Jing, Zonglei, Xiao, Yisong, Zou, Quanchen, Liu, Aishan, Liang, Siyuan, Zhang, Xiangzheng, Liu, Xianglong, Tao, Dacheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reasoning-Oriented Programming: Chaining Semantic Gadgets to Jailbreak Large Vision Language Models
von: Zou, Quanchen, et al.
Veröffentlicht: (2026)
von: Zou, Quanchen, et al.
Veröffentlicht: (2026)
Unveiling the Safety of GPT-4o: An Empirical Study using Jailbreak Attacks
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
PRISM: Programmatic Reasoning with Image Sequence Manipulation for LVLM Jailbreaking
von: Zou, Quanchen, et al.
Veröffentlicht: (2025)
von: Zou, Quanchen, et al.
Veröffentlicht: (2025)
SPARK: Jailbreaking T2V Models by Synergistically Prompting Auditory and Recontextualized Knowledge
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs
von: Xu, Wenzhuo, et al.
Veröffentlicht: (2026)
von: Xu, Wenzhuo, et al.
Veröffentlicht: (2026)
Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?
von: Mu, Junjie, et al.
Veröffentlicht: (2025)
von: Mu, Junjie, et al.
Veröffentlicht: (2025)
Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
Probabilistic Modeling of Jailbreak on Multimodal LLMs: From Quantification to Application
von: Xu, Wenzhuo, et al.
Veröffentlicht: (2025)
von: Xu, Wenzhuo, et al.
Veröffentlicht: (2025)
SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety
von: Liu, Zhe, et al.
Veröffentlicht: (2026)
von: Liu, Zhe, et al.
Veröffentlicht: (2026)
Evolving Deception: When Agents Evolve, Deception Wins
von: Ying, Zonghao, et al.
Veröffentlicht: (2026)
von: Ying, Zonghao, et al.
Veröffentlicht: (2026)
Two Frames Matter: A Temporal Attack for Text-to-Video Model Jailbreaking
von: Chen, Moyang, et al.
Veröffentlicht: (2026)
von: Chen, Moyang, et al.
Veröffentlicht: (2026)
SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
Robust Privacy: Inference-Time Privacy through Certified Robustness
von: Jin, Jiankai, et al.
Veröffentlicht: (2026)
von: Jin, Jiankai, et al.
Veröffentlicht: (2026)
AgentVisor: Defending LLM Agents Against Prompt Injection via Semantic Virtualization
von: Ying, Zonghao, et al.
Veröffentlicht: (2026)
von: Ying, Zonghao, et al.
Veröffentlicht: (2026)
Sequential Comics for Jailbreaking Multimodal Large Language Models via Structured Visual Storytelling
von: Zhang, Deyue, et al.
Veröffentlicht: (2025)
von: Zhang, Deyue, et al.
Veröffentlicht: (2025)
ME: Trigger Element Combination Backdoor Attack on Copyright Infringement
von: Yang, Feiyu, et al.
Veröffentlicht: (2025)
von: Yang, Feiyu, et al.
Veröffentlicht: (2025)
ICLShield: Exploring and Mitigating In-Context Learning Backdoor Attacks
von: Ren, Zhiyao, et al.
Veröffentlicht: (2025)
von: Ren, Zhiyao, et al.
Veröffentlicht: (2025)
T2VShield: Model-Agnostic Jailbreak Defense for Text-to-Video Models
von: Liang, Siyuan, et al.
Veröffentlicht: (2025)
von: Liang, Siyuan, et al.
Veröffentlicht: (2025)
AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
Reading Between the Pixels: An Inscriptive Jailbreak Attack on Text-to-Image Models
von: Ying, Zonghao, et al.
Veröffentlicht: (2026)
von: Ying, Zonghao, et al.
Veröffentlicht: (2026)
Compromising Embodied Agents with Contextual Backdoor Attacks
von: Liu, Aishan, et al.
Veröffentlicht: (2024)
von: Liu, Aishan, et al.
Veröffentlicht: (2024)
ELBA-Bench: An Efficient Learning Backdoor Attacks Benchmark for Large Language Models
von: Liu, Xuxu, et al.
Veröffentlicht: (2025)
von: Liu, Xuxu, et al.
Veröffentlicht: (2025)
CogMorph: Cognitive Morphing Attacks for Text-to-Image Models
von: Jing, Zonglei, et al.
Veröffentlicht: (2025)
von: Jing, Zonglei, et al.
Veröffentlicht: (2025)
Demystifying and Detecting Agentic Workflow Injection Vulnerabilities in GitHub Actions
von: Wang, Shenao, et al.
Veröffentlicht: (2026)
von: Wang, Shenao, et al.
Veröffentlicht: (2026)
Uncovering Security Threats and Architecting Defenses in Autonomous Agents: A Case Study of OpenClaw
von: Ying, Zonghao, et al.
Veröffentlicht: (2026)
von: Ying, Zonghao, et al.
Veröffentlicht: (2026)
MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks
von: Zhang, Xinkai, et al.
Veröffentlicht: (2026)
von: Zhang, Xinkai, et al.
Veröffentlicht: (2026)
SafeMobile: Chain-level Jailbreak Detection and Automated Evaluation for Multimodal Mobile Agents
von: Liang, Siyuan, et al.
Veröffentlicht: (2025)
von: Liang, Siyuan, et al.
Veröffentlicht: (2025)
Pandora: Jailbreak GPTs by Retrieval Augmented Generation Poisoning
von: Deng, Gelei, et al.
Veröffentlicht: (2024)
von: Deng, Gelei, et al.
Veröffentlicht: (2024)
Detoxifying Large Language Models via Autoregressive Reward Guided Representation Editing
von: Xiao, Yisong, et al.
Veröffentlicht: (2025)
von: Xiao, Yisong, et al.
Veröffentlicht: (2025)
SoK: Understanding Vulnerabilities in the Large Language Model Supply Chain
von: Wang, Shenao, et al.
Veröffentlicht: (2025)
von: Wang, Shenao, et al.
Veröffentlicht: (2025)
Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks
von: Gibbs, Tom, et al.
Veröffentlicht: (2024)
von: Gibbs, Tom, et al.
Veröffentlicht: (2024)
PolyJailbreak: Cross-Modal Jailbreaking Attacks on Black-Box Multimodal LLMs
von: Wang, Xinkai, et al.
Veröffentlicht: (2025)
von: Wang, Xinkai, et al.
Veröffentlicht: (2025)
ICON: Intent-Context Coupling for Efficient Multi-Turn Jailbreak Attack
von: Lin, Xingwei, et al.
Veröffentlicht: (2026)
von: Lin, Xingwei, et al.
Veröffentlicht: (2026)
The Trojan Example: Jailbreaking LLMs through Template Filling and Unsafety Reasoning
von: Liu, Mingrui, et al.
Veröffentlicht: (2025)
von: Liu, Mingrui, et al.
Veröffentlicht: (2025)
DLP: towards active defense against backdoor attacks with decoupled learning process
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
NBA: defensive distillation for backdoor removal via neural behavior alignment
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
Let the Bees Find the Weak Spots: A Path Planning Perspective on Multi-Turn Jailbreak Attacks against LLMs
von: Liu, Yize, et al.
Veröffentlicht: (2025)
von: Liu, Yize, et al.
Veröffentlicht: (2025)
External Data Extraction Attacks against Retrieval-Augmented Large Language Models
von: He, Yu, et al.
Veröffentlicht: (2025)
von: He, Yu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Reasoning-Oriented Programming: Chaining Semantic Gadgets to Jailbreak Large Vision Language Models
von: Zou, Quanchen, et al.
Veröffentlicht: (2026) -
Unveiling the Safety of GPT-4o: An Empirical Study using Jailbreak Attacks
von: Ying, Zonghao, et al.
Veröffentlicht: (2024) -
PRISM: Programmatic Reasoning with Image Sequence Manipulation for LVLM Jailbreaking
von: Zou, Quanchen, et al.
Veröffentlicht: (2025) -
SPARK: Jailbreaking T2V Models by Synergistically Prompting Auditory and Recontextualized Knowledge
von: Ying, Zonghao, et al.
Veröffentlicht: (2025) -
DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs
von: Xu, Wenzhuo, et al.
Veröffentlicht: (2026)