LPG: Balancing Efficiency and Policy Reasoning in Latent Policy Guardrails
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Nanxi, Zhao, Zhengyue, Xiao, Chaowei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PRISM: Robust VLM Alignment with Principled Reasoning for Integrated Safety in Multimodality
von: Li, Nanxi, et al.
Veröffentlicht: (2025)
von: Li, Nanxi, et al.
Veröffentlicht: (2025)
SafeVision: Efficient Image Guardrail with Robust Policy Adherence and Explainability
von: Xu, Peiyang, et al.
Veröffentlicht: (2025)
von: Xu, Peiyang, et al.
Veröffentlicht: (2025)
AutoDAN-Reasoning: Enhancing Strategies Exploration based Jailbreak Attacks with Test-Time Scaling
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2025)
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2025)
Privacy Policy Enforcement Guardrails for Data-Sensitive Retrieval-Augmented Generation
von: Zafar, Osama, et al.
Veröffentlicht: (2026)
von: Zafar, Osama, et al.
Veröffentlicht: (2026)
No Free Lunch with Guardrails
von: Kumar, Divyanshu, et al.
Veröffentlicht: (2025)
von: Kumar, Divyanshu, et al.
Veröffentlicht: (2025)
Provably Secure Agent Guardrail
von: Wu, Benlong, et al.
Veröffentlicht: (2026)
von: Wu, Benlong, et al.
Veröffentlicht: (2026)
ReasoningBomb: A Stealthy Denial-of-Service Attack by Inducing Pathologically Long Reasoning in Large Reasoning Models
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2026)
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2026)
OET: Optimization-based prompt injection Evaluation Toolkit
von: Pan, Jinsheng, et al.
Veröffentlicht: (2025)
von: Pan, Jinsheng, et al.
Veröffentlicht: (2025)
RePD: Defending Jailbreak Attack through a Retrieval-based Prompt Decomposition Process
von: Wang, Peiran, et al.
Veröffentlicht: (2024)
von: Wang, Peiran, et al.
Veröffentlicht: (2024)
A Comparative Evaluation of AI Agent Security Guardrails
von: Li, Qi, et al.
Veröffentlicht: (2026)
von: Li, Qi, et al.
Veröffentlicht: (2026)
AgentSys: Secure and Dynamic LLM Agents Through Explicit Hierarchical Memory Management
von: Wen, Ruoyao, et al.
Veröffentlicht: (2026)
von: Wen, Ruoyao, et al.
Veröffentlicht: (2026)
Enhancing Guardrails for Safe and Secure Healthcare AI
von: Gangavarapu, Ananya
Veröffentlicht: (2024)
von: Gangavarapu, Ananya
Veröffentlicht: (2024)
SoK: Evaluating Jailbreak Guardrails for Large Language Models
von: Wang, Xunguang, et al.
Veröffentlicht: (2025)
von: Wang, Xunguang, et al.
Veröffentlicht: (2025)
ARMOR: Aligning Secure and Safe Large Language Models via Meticulous Reasoning
von: Zhao, Zhengyue, et al.
Veröffentlicht: (2025)
von: Zhao, Zhengyue, et al.
Veröffentlicht: (2025)
Defenses Against Prompt Attacks Learn Surface Heuristics
von: Li, Shawn, et al.
Veröffentlicht: (2026)
von: Li, Shawn, et al.
Veröffentlicht: (2026)
JailDAM: Jailbreak Detection with Adaptive Memory for Vision-Language Model
von: Nian, Yi, et al.
Veröffentlicht: (2025)
von: Nian, Yi, et al.
Veröffentlicht: (2025)
Doxing via the Lens: Revealing Location-related Privacy Leakage on Multi-modal Large Reasoning Models
von: Luo, Weidi, et al.
Veröffentlicht: (2025)
von: Luo, Weidi, et al.
Veröffentlicht: (2025)
WIPI: A New Web Threat for LLM-Driven Web Agents
von: Wu, Fangzhou, et al.
Veröffentlicht: (2024)
von: Wu, Fangzhou, et al.
Veröffentlicht: (2024)
AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting
von: Wang, Yu, et al.
Veröffentlicht: (2024)
von: Wang, Yu, et al.
Veröffentlicht: (2024)
ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack
von: Li, Hao, et al.
Veröffentlicht: (2026)
von: Li, Hao, et al.
Veröffentlicht: (2026)
Cognitive Cybersecurity for Artificial Intelligence: Guardrail Engineering with CCS-7
von: Aydin, Yuksel
Veröffentlicht: (2025)
von: Aydin, Yuksel
Veröffentlicht: (2025)
Breaking Guardrails, Facing Walls: Insights on Adversarial AI for Defenders & Researchers
von: Bertollo, Giacomo, et al.
Veröffentlicht: (2025)
von: Bertollo, Giacomo, et al.
Veröffentlicht: (2025)
Balancing Privacy and Efficiency: Music Information Retrieval via Additive Homomorphic Encryption
von: Wang, William Zerong, et al.
Veröffentlicht: (2025)
von: Wang, William Zerong, et al.
Veröffentlicht: (2025)
LLM Access Shield: Domain-Specific LLM Framework for Privacy Policy Compliance
von: Wang, Yu, et al.
Veröffentlicht: (2025)
von: Wang, Yu, et al.
Veröffentlicht: (2025)
DRIFT: Dynamic Rule-Based Defense with Injection Isolation for Securing LLM Agents
von: Li, Hao, et al.
Veröffentlicht: (2025)
von: Li, Hao, et al.
Veröffentlicht: (2025)
Active Honeypot Guardrail System: Probing and Confirming Multi-Turn LLM Jailbreaks
von: Wu, ChenYu, et al.
Veröffentlicht: (2025)
von: Wu, ChenYu, et al.
Veröffentlicht: (2025)
JPU: Bridging Jailbreak Defense and Unlearning via On-Policy Path Rectification
von: Wang, Xi, et al.
Veröffentlicht: (2026)
von: Wang, Xi, et al.
Veröffentlicht: (2026)
A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems
von: Wu, Fangzhou, et al.
Veröffentlicht: (2024)
von: Wu, Fangzhou, et al.
Veröffentlicht: (2024)
On Automating Security Policies with Contemporary LLMs
von: Saura, Pablo Fernández, et al.
Veröffentlicht: (2025)
von: Saura, Pablo Fernández, et al.
Veröffentlicht: (2025)
OneShield -- the Next Generation of LLM Guardrails
von: DeLuca, Chad, et al.
Veröffentlicht: (2025)
von: DeLuca, Chad, et al.
Veröffentlicht: (2025)
How Good LLM-Generated Password Policies Are?
von: Vaidya, Vivek, et al.
Veröffentlicht: (2025)
von: Vaidya, Vivek, et al.
Veröffentlicht: (2025)
RAGent: Retrieval-based Access Control Policy Generation
von: Jayasundara, Sakuna Harinda, et al.
Veröffentlicht: (2024)
von: Jayasundara, Sakuna Harinda, et al.
Veröffentlicht: (2024)
Architecting Secure AI Agents: Perspectives on System-Level Defenses Against Indirect Prompt Injection Attacks
von: Xiang, Chong, et al.
Veröffentlicht: (2026)
von: Xiang, Chong, et al.
Veröffentlicht: (2026)
InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models
von: Li, Hao, et al.
Veröffentlicht: (2024)
von: Li, Hao, et al.
Veröffentlicht: (2024)
Improving LLM Reasoning for Vulnerability Detection via Group Relative Policy Optimization
von: Simoni, Marco, et al.
Veröffentlicht: (2025)
von: Simoni, Marco, et al.
Veröffentlicht: (2025)
Proof-of-Guardrail in AI Agents and What (Not) to Trust from It
von: Jin, Xisen, et al.
Veröffentlicht: (2026)
von: Jin, Xisen, et al.
Veröffentlicht: (2026)
NeuroFilter: Privacy Guardrails for Conversational LLM Agents
von: Das, Saswat, et al.
Veröffentlicht: (2026)
von: Das, Saswat, et al.
Veröffentlicht: (2026)
APEX: Agent Payment Execution with Policy for Autonomous Agent API Access
von: Uddin, Mohd Safwan, et al.
Veröffentlicht: (2026)
von: Uddin, Mohd Safwan, et al.
Veröffentlicht: (2026)
AudAgent: Automated Auditing of Privacy Policy Compliance in AI Agents
von: Zheng, Ye, et al.
Veröffentlicht: (2025)
von: Zheng, Ye, et al.
Veröffentlicht: (2025)
A BERT-based Empirical Study of Privacy Policies' Compliance with GDPR
von: Zhang, Lu, et al.
Veröffentlicht: (2024)
von: Zhang, Lu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
PRISM: Robust VLM Alignment with Principled Reasoning for Integrated Safety in Multimodality
von: Li, Nanxi, et al.
Veröffentlicht: (2025) -
SafeVision: Efficient Image Guardrail with Robust Policy Adherence and Explainability
von: Xu, Peiyang, et al.
Veröffentlicht: (2025) -
AutoDAN-Reasoning: Enhancing Strategies Exploration based Jailbreak Attacks with Test-Time Scaling
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2025) -
Privacy Policy Enforcement Guardrails for Data-Sensitive Retrieval-Augmented Generation
von: Zafar, Osama, et al.
Veröffentlicht: (2026) -
No Free Lunch with Guardrails
von: Kumar, Divyanshu, et al.
Veröffentlicht: (2025)