X-Boundary: Establishing Exact Safety Boundary to Shield LLMs from Multi-Turn Jailbreaks without Compromising Usability
Fuente:
arXiv
Salvato in:
| Autori principali: | Lu, Xiaoya, Liu, Dongrui, Yu, Yi, Xu, Luxin, Shao, Jing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LLMs Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions
di: Hu, Xuhao, et al.
Pubblicazione: (2025)
di: Hu, Xuhao, et al.
Pubblicazione: (2025)
Overlooked Safety Vulnerability in LLMs: Malicious Intelligent Optimization Algorithm Request and its Jailbreak
di: Gu, Haoran, et al.
Pubblicazione: (2026)
di: Gu, Haoran, et al.
Pubblicazione: (2026)
Let the Bees Find the Weak Spots: A Path Planning Perspective on Multi-Turn Jailbreak Attacks against LLMs
di: Liu, Yize, et al.
Pubblicazione: (2025)
di: Liu, Yize, et al.
Pubblicazione: (2025)
AgentShield: Deception-based Compromise Detection for Tool-using LLM Agents
di: Rassul, Yassin H., et al.
Pubblicazione: (2026)
di: Rassul, Yassin H., et al.
Pubblicazione: (2026)
ShieldLearner: A New Paradigm for Jailbreak Attack Defense in LLMs
di: Ni, Ziyi, et al.
Pubblicazione: (2025)
di: Ni, Ziyi, et al.
Pubblicazione: (2025)
A Simple and Efficient Jailbreak Method Exploiting LLMs' Helpfulness
di: Luo, Xuan, et al.
Pubblicazione: (2025)
di: Luo, Xuan, et al.
Pubblicazione: (2025)
Token-Level Constraint Boundary Search for Jailbreaking Text-to-Image Models
di: Liu, Jiangtao, et al.
Pubblicazione: (2025)
di: Liu, Jiangtao, et al.
Pubblicazione: (2025)
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
di: Zhang, Chiyu, et al.
Pubblicazione: (2025)
di: Zhang, Chiyu, et al.
Pubblicazione: (2025)
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient Analysis
di: Xie, Yueqi, et al.
Pubblicazione: (2024)
di: Xie, Yueqi, et al.
Pubblicazione: (2024)
BAIT: Boundary-Guided Disclosure Escalation via Self-Conditioned Reasoning
di: Luo, Xuan, et al.
Pubblicazione: (2026)
di: Luo, Xuan, et al.
Pubblicazione: (2026)
Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks
di: Gibbs, Tom, et al.
Pubblicazione: (2024)
di: Gibbs, Tom, et al.
Pubblicazione: (2024)
LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
di: Li, Nathaniel, et al.
Pubblicazione: (2024)
di: Li, Nathaniel, et al.
Pubblicazione: (2024)
Siren: A Learning-Based Multi-Turn Attack Framework for Simulating Real-World Human Jailbreak Behaviors
di: Zhao, Yi, et al.
Pubblicazione: (2025)
di: Zhao, Yi, et al.
Pubblicazione: (2025)
One Trigger Token Is Enough: A Defense Strategy for Balancing Safety and Usability in Large Language Models
di: Gu, Haoran, et al.
Pubblicazione: (2025)
di: Gu, Haoran, et al.
Pubblicazione: (2025)
VLSBench: Unveiling Visual Leakage in Multimodal Safety
di: Hu, Xuhao, et al.
Pubblicazione: (2024)
di: Hu, Xuhao, et al.
Pubblicazione: (2024)
Model X-ray:Detecting Backdoored Models via Decision Boundary
di: Su, Yanghao, et al.
Pubblicazione: (2024)
di: Su, Yanghao, et al.
Pubblicazione: (2024)
AlienLM: Alienization of Language for API-Boundary Privacy in Black-Box LLMs
di: Kim, Jaehee, et al.
Pubblicazione: (2026)
di: Kim, Jaehee, et al.
Pubblicazione: (2026)
Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context Injection
di: Miao, Ziqi, et al.
Pubblicazione: (2025)
di: Miao, Ziqi, et al.
Pubblicazione: (2025)
Mitigating Jailbreaks with Intent-Aware LLMs
di: Yeo, Wei Jie, et al.
Pubblicazione: (2025)
di: Yeo, Wei Jie, et al.
Pubblicazione: (2025)
X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents
di: Rahman, Salman, et al.
Pubblicazione: (2025)
di: Rahman, Salman, et al.
Pubblicazione: (2025)
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
di: Shen, Guobin, et al.
Pubblicazione: (2025)
di: Shen, Guobin, et al.
Pubblicazione: (2025)
Soft Begging: Modular and Efficient Shielding of LLMs against Prompt Injection and Jailbreaking based on Prompt Tuning
di: Ostermann, Simon, et al.
Pubblicazione: (2024)
di: Ostermann, Simon, et al.
Pubblicazione: (2024)
Tempest: Autonomous Multi-Turn Jailbreaking of Large Language Models with Tree Search
di: Zhou, Andy, et al.
Pubblicazione: (2025)
di: Zhou, Andy, et al.
Pubblicazione: (2025)
Jailbreaking LLMs via Semantically Relevant Nested Scenarios with Targeted Toxic Knowledge
di: Xu, Ning, et al.
Pubblicazione: (2025)
di: Xu, Ning, et al.
Pubblicazione: (2025)
AI Safety vs. AI Security: Demystifying the Distinction and Boundaries
di: Lin, Zhiqiang, et al.
Pubblicazione: (2025)
di: Lin, Zhiqiang, et al.
Pubblicazione: (2025)
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
di: Hu, Xiaomeng, et al.
Pubblicazione: (2025)
di: Hu, Xiaomeng, et al.
Pubblicazione: (2025)
Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models
di: Nihal, Ragib Amin, et al.
Pubblicazione: (2025)
di: Nihal, Ragib Amin, et al.
Pubblicazione: (2025)
RED QUEEN: Safeguarding Large Language Models against Concealed Multi-Turn Jailbreaking
di: Jiang, Yifan, et al.
Pubblicazione: (2024)
di: Jiang, Yifan, et al.
Pubblicazione: (2024)
One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMs
di: Li, Linbao, et al.
Pubblicazione: (2025)
di: Li, Linbao, et al.
Pubblicazione: (2025)
Jailbreak Antidote: Runtime Safety-Utility Balance via Sparse Representation Adjustment in Large Language Models
di: Shen, Guobin, et al.
Pubblicazione: (2024)
di: Shen, Guobin, et al.
Pubblicazione: (2024)
Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs
di: Xu, Zhao, et al.
Pubblicazione: (2024)
di: Xu, Zhao, et al.
Pubblicazione: (2024)
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
di: Chen, Yunhao, et al.
Pubblicazione: (2025)
di: Chen, Yunhao, et al.
Pubblicazione: (2025)
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
di: Wang, Jiongxiao, et al.
Pubblicazione: (2024)
di: Wang, Jiongxiao, et al.
Pubblicazione: (2024)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
di: Liu, Fan, et al.
Pubblicazione: (2024)
di: Liu, Fan, et al.
Pubblicazione: (2024)
Jailbreaking LLMs via Calibration
di: Lu, Yuxuan, et al.
Pubblicazione: (2026)
di: Lu, Yuxuan, et al.
Pubblicazione: (2026)
Jailbreak Distillation: Renewable Safety Benchmarking
di: Zhang, Jingyu, et al.
Pubblicazione: (2025)
di: Zhang, Jingyu, et al.
Pubblicazione: (2025)
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
di: Zhou, Zhenhong, et al.
Pubblicazione: (2024)
di: Zhou, Zhenhong, et al.
Pubblicazione: (2024)
SafeAligner: Safety Alignment against Jailbreak Attacks via Response Disparity Guidance
di: Huang, Caishuang, et al.
Pubblicazione: (2024)
di: Huang, Caishuang, et al.
Pubblicazione: (2024)
Documenti analoghi
-
LLMs Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions
di: Hu, Xuhao, et al.
Pubblicazione: (2025) -
Overlooked Safety Vulnerability in LLMs: Malicious Intelligent Optimization Algorithm Request and its Jailbreak
di: Gu, Haoran, et al.
Pubblicazione: (2026) -
Let the Bees Find the Weak Spots: A Path Planning Perspective on Multi-Turn Jailbreak Attacks against LLMs
di: Liu, Yize, et al.
Pubblicazione: (2025) -
AgentShield: Deception-based Compromise Detection for Tool-using LLM Agents
di: Rassul, Yassin H., et al.
Pubblicazione: (2026) -
ShieldLearner: A New Paradigm for Jailbreak Attack Defense in LLMs
di: Ni, Ziyi, et al.
Pubblicazione: (2025)