Provably Secure Agent Guardrail
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Benlong, Zhang, Weiming, Chen, Kejiang, Fang, Han, Yu, Nenghai |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AutoPT: How Far Are We from the End2End Automated Web Penetration Testing?
by: Wu, Benlong, et al.
Published: (2024)
by: Wu, Benlong, et al.
Published: (2024)
STEAD: Robust Provably Secure Linguistic Steganography with Diffusion Language Model
by: Qi, Yuang, et al.
Published: (2026)
by: Qi, Yuang, et al.
Published: (2026)
A high-capacity linguistic steganography based on entropy-driven rank-token mapping
by: Jiang, Jun, et al.
Published: (2025)
by: Jiang, Jun, et al.
Published: (2025)
Provably Secure Disambiguating Neural Linguistic Steganography
by: Qi, Yuang, et al.
Published: (2024)
by: Qi, Yuang, et al.
Published: (2024)
Provably Secure Public-Key Steganography Based on Admissible Encoding
by: Zhang, Xin, et al.
Published: (2025)
by: Zhang, Xin, et al.
Published: (2025)
Gaussian Shading: Provable Performance-Lossless Image Watermarking for Diffusion Models
by: Yang, Zijin, et al.
Published: (2024)
by: Yang, Zijin, et al.
Published: (2024)
A Comparative Evaluation of AI Agent Security Guardrails
by: Li, Qi, et al.
Published: (2026)
by: Li, Qi, et al.
Published: (2026)
De-AntiFake: Rethinking the Protective Perturbations Against Voice Cloning Attacks
by: Fan, Wei, et al.
Published: (2025)
by: Fan, Wei, et al.
Published: (2025)
WavInWav: Time-domain Speech Hiding via Invertible Neural Network
by: Fan, Wei, et al.
Published: (2025)
by: Fan, Wei, et al.
Published: (2025)
SQL Injection Jailbreak: A Structural Disaster of Large Language Models
by: Zhao, Jiawei, et al.
Published: (2024)
by: Zhao, Jiawei, et al.
Published: (2024)
Performance-lossless Black-box Model Watermarking
by: Zhao, Na, et al.
Published: (2023)
by: Zhao, Na, et al.
Published: (2023)
Membership Inference Attacks on Tokenizers of Large Language Models
by: Tong, Meng, et al.
Published: (2025)
by: Tong, Meng, et al.
Published: (2025)
GIFDL: Generated Image Fluctuation Distortion Learning for Enhancing Steganographic Security
by: Wang, Xiangkun, et al.
Published: (2025)
by: Wang, Xiangkun, et al.
Published: (2025)
Enhancing Guardrails for Safe and Secure Healthcare AI
by: Gangavarapu, Ananya
Published: (2024)
by: Gangavarapu, Ananya
Published: (2024)
State-Dependent Safety Failures in Multi-Turn Language Model Interaction
by: Li, Pengcheng, et al.
Published: (2026)
by: Li, Pengcheng, et al.
Published: (2026)
Prefix Guidance: A Steering Wheel for Large Language Models to Defend Against Jailbreak Attacks
by: Zhao, Jiawei, et al.
Published: (2024)
by: Zhao, Jiawei, et al.
Published: (2024)
PSRT: Accelerating LRM-based Guard Models via Prefilled Safe Reasoning Traces
by: Zhao, Jiawei, et al.
Published: (2025)
by: Zhao, Jiawei, et al.
Published: (2025)
Active Honeypot Guardrail System: Probing and Confirming Multi-Turn LLM Jailbreaks
by: Wu, ChenYu, et al.
Published: (2025)
by: Wu, ChenYu, et al.
Published: (2025)
Provably Secure Retrieval-Augmented Generation
by: Zhou, Pengcheng, et al.
Published: (2025)
by: Zhou, Pengcheng, et al.
Published: (2025)
Character as a Latent Variable in Large Language Models: A Mechanistic Account of Emergent Misalignment and Conditional Safety Failures
by: Su, Yanghao, et al.
Published: (2026)
by: Su, Yanghao, et al.
Published: (2026)
Towards Provable (In)Secure Model Weight Release Schemes
by: Yang, Xin, et al.
Published: (2025)
by: Yang, Xin, et al.
Published: (2025)
No Free Lunch with Guardrails
by: Kumar, Divyanshu, et al.
Published: (2025)
by: Kumar, Divyanshu, et al.
Published: (2025)
Symbolic Guardrails for Domain-Specific Agents: Stronger Safety and Security Guarantees Without Sacrificing Utility
by: Hong, Yining, et al.
Published: (2026)
by: Hong, Yining, et al.
Published: (2026)
Silent Guardian: Protecting Text from Malicious Exploitation by Large Language Models
by: Zhao, Jiawei, et al.
Published: (2023)
by: Zhao, Jiawei, et al.
Published: (2023)
SparSamp: Efficient Provably Secure Steganography Based on Sparse Sampling
by: Wang, Yaofei, et al.
Published: (2025)
by: Wang, Yaofei, et al.
Published: (2025)
Training with Differential Privacy: A Gradient-Preserving Noise Reduction Approach with Provable Security
by: Wang, Haodi, et al.
Published: (2024)
by: Wang, Haodi, et al.
Published: (2024)
LLM Agents Should Employ Security Principles
by: Zhang, Kaiyuan, et al.
Published: (2025)
by: Zhang, Kaiyuan, et al.
Published: (2025)
Proof-of-Guardrail in AI Agents and What (Not) to Trust from It
by: Jin, Xisen, et al.
Published: (2026)
by: Jin, Xisen, et al.
Published: (2026)
SoK: Evaluating Jailbreak Guardrails for Large Language Models
by: Wang, Xunguang, et al.
Published: (2025)
by: Wang, Xunguang, et al.
Published: (2025)
SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety
by: Liu, Zhe, et al.
Published: (2026)
by: Liu, Zhe, et al.
Published: (2026)
NeuroFilter: Privacy Guardrails for Conversational LLM Agents
by: Das, Saswat, et al.
Published: (2026)
by: Das, Saswat, et al.
Published: (2026)
Security of AI Agents
by: He, Yifeng, et al.
Published: (2024)
by: He, Yifeng, et al.
Published: (2024)
AgentWard: A Lifecycle Security Architecture for Autonomous AI Agents
by: Zhang, Yixiang, et al.
Published: (2026)
by: Zhang, Yixiang, et al.
Published: (2026)
On the Vulnerability of Text Sanitization
by: Tong, Meng, et al.
Published: (2024)
by: Tong, Meng, et al.
Published: (2024)
Turning Your Strength into Watermark: Watermarking Large Language Model via Knowledge Injection
by: Li, Shuai, et al.
Published: (2023)
by: Li, Shuai, et al.
Published: (2023)
Towards Secure Agent Skills: Architecture, Threat Taxonomy, and Security Analysis
by: Li, Zhiyuan, et al.
Published: (2026)
by: Li, Zhiyuan, et al.
Published: (2026)
Provably Secure Robust Image Steganography via Cross-Modal Error Correction
by: Qi, Yuang, et al.
Published: (2024)
by: Qi, Yuang, et al.
Published: (2024)
LPG: Balancing Efficiency and Policy Reasoning in Latent Policy Guardrails
by: Li, Nanxi, et al.
Published: (2026)
by: Li, Nanxi, et al.
Published: (2026)
Cognitive Cybersecurity for Artificial Intelligence: Guardrail Engineering with CCS-7
by: Aydin, Yuksel
Published: (2025)
by: Aydin, Yuksel
Published: (2025)
MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents
by: Zhu, Kaijie, et al.
Published: (2025)
by: Zhu, Kaijie, et al.
Published: (2025)
Similar Items
-
AutoPT: How Far Are We from the End2End Automated Web Penetration Testing?
by: Wu, Benlong, et al.
Published: (2024) -
STEAD: Robust Provably Secure Linguistic Steganography with Diffusion Language Model
by: Qi, Yuang, et al.
Published: (2026) -
A high-capacity linguistic steganography based on entropy-driven rank-token mapping
by: Jiang, Jun, et al.
Published: (2025) -
Provably Secure Disambiguating Neural Linguistic Steganography
by: Qi, Yuang, et al.
Published: (2024) -
Provably Secure Public-Key Steganography Based on Admissible Encoding
by: Zhang, Xin, et al.
Published: (2025)