Bag of Tricks for Subverting Reasoning-based Safety Guardrails
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Shuo, Han, Zhen, Chen, Haokun, He, Bailan, Si, Shengyun, Wu, Jingpei, Torr, Philip, Tresp, Volker, Gu, Jindong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Deep Research Brings Deeper Harm
by: Chen, Shuo, et al.
Published: (2025)
by: Chen, Shuo, et al.
Published: (2025)
Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks?
by: Chen, Shuo, et al.
Published: (2024)
by: Chen, Shuo, et al.
Published: (2024)
Stop Reasoning! When Multimodal LLM with Chain-of-Thought Reasoning Meets Adversarial Image
by: Wang, Zefeng, et al.
Published: (2024)
by: Wang, Zefeng, et al.
Published: (2024)
Reimagining Safety Alignment with An Image
by: Xia, Yifan, et al.
Published: (2025)
by: Xia, Yifan, et al.
Published: (2025)
Multimodal Pragmatic Jailbreak on Text-to-image Models
by: Liu, Tong, et al.
Published: (2024)
by: Liu, Tong, et al.
Published: (2024)
Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs
by: Xu, Zhao, et al.
Published: (2024)
by: Xu, Zhao, et al.
Published: (2024)
GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio
by: Zhu, Zhenhao, et al.
Published: (2026)
by: Zhu, Zhenhao, et al.
Published: (2026)
Correcting Subverted Random Oracles
by: Russell, Alexander, et al.
Published: (2024)
by: Russell, Alexander, et al.
Published: (2024)
LLM Jailbreak Detection for (Almost) Free!
by: Chen, Guorui, et al.
Published: (2025)
by: Chen, Guorui, et al.
Published: (2025)
PoseGuard: Pose-Guided Generation with Safety Guardrails
by: Wang, Kongxin, et al.
Published: (2025)
by: Wang, Kongxin, et al.
Published: (2025)
Poly-Guard: Massive Multi-Domain Safety Policy-Grounded Guardrail Dataset
by: Kang, Mintong, et al.
Published: (2025)
by: Kang, Mintong, et al.
Published: (2025)
Provably Secure Agent Guardrail
by: Wu, Benlong, et al.
Published: (2026)
by: Wu, Benlong, et al.
Published: (2026)
SkillJect: Effectively Automating Skill-Based Prompt Injection for Skill-Enabled Agents
by: Jia, Xiaojun, et al.
Published: (2026)
by: Jia, Xiaojun, et al.
Published: (2026)
ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models
by: Zhao, Yunhan, et al.
Published: (2026)
by: Zhao, Yunhan, et al.
Published: (2026)
Black-Box Guardrail Reverse-engineering Attack
by: Yao, Hongwei, et al.
Published: (2025)
by: Yao, Hongwei, et al.
Published: (2025)
When AIOps Become "AI Oops": Subverting LLM-driven IT Operations via Telemetry Manipulation
by: Pasquini, Dario, et al.
Published: (2025)
by: Pasquini, Dario, et al.
Published: (2025)
LPG: Balancing Efficiency and Policy Reasoning in Latent Policy Guardrails
by: Li, Nanxi, et al.
Published: (2026)
by: Li, Nanxi, et al.
Published: (2026)
Imprompter: Tricking LLM Agents into Improper Tool Use
by: Fu, Xiaohan, et al.
Published: (2024)
by: Fu, Xiaohan, et al.
Published: (2024)
Triaging Threats to Specialized Guardrails
by: Mo, Wenjie Jacky, et al.
Published: (2026)
by: Mo, Wenjie Jacky, et al.
Published: (2026)
Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images
by: Gao, Kuofeng, et al.
Published: (2024)
by: Gao, Kuofeng, et al.
Published: (2024)
TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts
by: Chu, Hua-Rong, et al.
Published: (2026)
by: Chu, Hua-Rong, et al.
Published: (2026)
GSPR: Aligning LLM Safeguards as Generalizable Safety Policy Reasoners
by: Li, Haoran, et al.
Published: (2025)
by: Li, Haoran, et al.
Published: (2025)
Simple Perturbations Subvert Ethereum Phishing Transactions Detection: An Empirical Analysis
by: Alghureid, Ahod, et al.
Published: (2024)
by: Alghureid, Ahod, et al.
Published: (2024)
SpatialJB: How Text Distribution Art Becomes the "Jailbreak Key" for LLM Guardrails
by: Mou, Zhiyi, et al.
Published: (2026)
by: Mou, Zhiyi, et al.
Published: (2026)
No Free Lunch with Guardrails
by: Kumar, Divyanshu, et al.
Published: (2025)
by: Kumar, Divyanshu, et al.
Published: (2025)
SGuard-v1: Safety Guardrail for Large Language Models
by: Lee, JoonHo, et al.
Published: (2025)
by: Lee, JoonHo, et al.
Published: (2025)
Evaluating the Robustness of Large Language Model Safety Guardrails Against Adversarial Attacks
by: Young, Richard J.
Published: (2025)
by: Young, Richard J.
Published: (2025)
ALMGuard: Safety Shortcuts and Where to Find Them as Guardrails for Audio-Language Models
by: Jin, Weifei, et al.
Published: (2025)
by: Jin, Weifei, et al.
Published: (2025)
Teaching an Old Dog New Tricks: Verifiable FHE Using Commodity Hardware
by: Drean, Jules, et al.
Published: (2024)
by: Drean, Jules, et al.
Published: (2024)
Peering Behind the Shield: Guardrail Identification in Large Language Models
by: Yang, Ziqing, et al.
Published: (2025)
by: Yang, Ziqing, et al.
Published: (2025)
Interpretable LLM Guardrails via Sparse Representation Steering
by: He, Zeqing, et al.
Published: (2025)
by: He, Zeqing, et al.
Published: (2025)
OpenGuardrails: A Configurable, Unified, and Scalable Guardrails Platform for Large Language Models
by: Wang, Thomas, et al.
Published: (2025)
by: Wang, Thomas, et al.
Published: (2025)
Benchmarking Safety Risks of Knowledge-Intensive Reasoning under Malicious Knowledge Editing
by: Mao, Qinghua, et al.
Published: (2026)
by: Mao, Qinghua, et al.
Published: (2026)
The Trojan Knowledge: Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Search
by: Wei, Rongzhe, et al.
Published: (2025)
by: Wei, Rongzhe, et al.
Published: (2025)
Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?
by: Chen, Shuo, et al.
Published: (2023)
by: Chen, Shuo, et al.
Published: (2023)
Active Honeypot Guardrail System: Probing and Confirming Multi-Turn LLM Jailbreaks
by: Wu, ChenYu, et al.
Published: (2025)
by: Wu, ChenYu, et al.
Published: (2025)
SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety
by: Liu, Zhe, et al.
Published: (2026)
by: Liu, Zhe, et al.
Published: (2026)
Auto-Tuning Safety Guardrails for Black-Box Large Language Models
by: Abdulkadir, Perry
Published: (2025)
by: Abdulkadir, Perry
Published: (2025)
Tricking LLM-Based NPCs into Spilling Secrets
by: Shiomi, Kyohei, et al.
Published: (2025)
by: Shiomi, Kyohei, et al.
Published: (2025)
Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets
by: Hsiung, Lei, et al.
Published: (2025)
by: Hsiung, Lei, et al.
Published: (2025)
Similar Items
-
Deep Research Brings Deeper Harm
by: Chen, Shuo, et al.
Published: (2025) -
Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks?
by: Chen, Shuo, et al.
Published: (2024) -
Stop Reasoning! When Multimodal LLM with Chain-of-Thought Reasoning Meets Adversarial Image
by: Wang, Zefeng, et al.
Published: (2024) -
Reimagining Safety Alignment with An Image
by: Xia, Yifan, et al.
Published: (2025) -
Multimodal Pragmatic Jailbreak on Text-to-image Models
by: Liu, Tong, et al.
Published: (2024)