ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wang, Yan, Chu, Zhixuan, Xue, Zihao, Bi, Zhen, Zhu, Bingyu, Chen, YueFeng, Yang, Zeyu, Lou, Jungang, Huang, Longtao, Zhang, Ningyu, Ren, Kui, Xue, Hui |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Make LLM Learn to Synthesize from Streaming Experiences through Feedback
par: Hu, Zhenlin, et autres
Publié: (2026)
par: Hu, Zhenlin, et autres
Publié: (2026)
Robust and Generalizable Safety Steering for Text-to-Image Diffusion Transformers
par: Xue, Zihao, et autres
Publié: (2026)
par: Xue, Zihao, et autres
Publié: (2026)
OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning
par: Zhu, Boyu, et autres
Publié: (2025)
par: Zhu, Boyu, et autres
Publié: (2025)
A Causal Explainable Guardrails for Large Language Models
par: Chu, Zhixuan, et autres
Publié: (2024)
par: Chu, Zhixuan, et autres
Publié: (2024)
Spatial Knowledge Graph-Guided Multimodal Synthesis
par: Xue, Yida, et autres
Publié: (2025)
par: Xue, Yida, et autres
Publié: (2025)
Poly-Guard: Massive Multi-Domain Safety Policy-Grounded Guardrail Dataset
par: Kang, Mintong, et autres
Publié: (2025)
par: Kang, Mintong, et autres
Publié: (2025)
Logical Structure as Knowledge: Enhancing LLM Reasoning via Structured Logical Knowledge Density Estimation
par: Bi, Zhen, et autres
Publié: (2025)
par: Bi, Zhen, et autres
Publié: (2025)
PoseGuard: Pose-Guided Generation with Safety Guardrails
par: Wang, Kongxin, et autres
Publié: (2025)
par: Wang, Kongxin, et autres
Publié: (2025)
ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models
par: Zhao, Yunhan, et autres
Publié: (2026)
par: Zhao, Yunhan, et autres
Publié: (2026)
Privacy Policy Enforcement Guardrails for Data-Sensitive Retrieval-Augmented Generation
par: Zafar, Osama, et autres
Publié: (2026)
par: Zafar, Osama, et autres
Publié: (2026)
MrGuard: A Multilingual Reasoning Guardrail for Universal LLM Safety
par: Yang, Yahan, et autres
Publié: (2025)
par: Yang, Yahan, et autres
Publié: (2025)
Thought Purity: A Defense Framework For Chain-of-Thought Attack
par: Xue, Zihao, et autres
Publié: (2025)
par: Xue, Zihao, et autres
Publié: (2025)
ConceptGuard: Neuro-Symbolic Safety Guardrails via Sparse Interpretable Jailbreak Concepts
par: Aswal, Darpan, et autres
Publié: (2025)
par: Aswal, Darpan, et autres
Publié: (2025)
Interpretable LLM Guardrails via Sparse Representation Steering
par: He, Zeqing, et autres
Publié: (2025)
par: He, Zeqing, et autres
Publié: (2025)
Prompt-Consistency Image Generation (PCIG): A Unified Framework Integrating LLMs, Knowledge Graphs, and Controllable Diffusion Models
par: Sun, Yichen, et autres
Publié: (2024)
par: Sun, Yichen, et autres
Publié: (2024)
Sora Detector: A Unified Hallucination Detection for Large Text-to-Video Models
par: Chu, Zhixuan, et autres
Publié: (2024)
par: Chu, Zhixuan, et autres
Publié: (2024)
UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models
par: Oh, Sejoon, et autres
Publié: (2024)
par: Oh, Sejoon, et autres
Publié: (2024)
SafeLens: Deliberate and Efficient Video Guardrails with Fast-and-Slow Screening
par: Nahin, Shahriar Kabir, et autres
Publié: (2026)
par: Nahin, Shahriar Kabir, et autres
Publié: (2026)
Why Steering Works: Toward a Unified View of Language Model Parameter Dynamics
par: Xu, Ziwen, et autres
Publié: (2026)
par: Xu, Ziwen, et autres
Publié: (2026)
Pivoting to Avoid Pitfalls: Trade Policy Uncertainty and Corporate ESG Performance
par: Xue Tan, et autres
Publié: (2025)
par: Xue Tan, et autres
Publié: (2025)
CodeGuard: Improving LLM Guardrails in CS Education
par: Raihan, Nishat, et autres
Publié: (2026)
par: Raihan, Nishat, et autres
Publié: (2026)
RegGuard: Legitimacy and Fairness Enforcement for Optimistic Rollups
par: Shang, Zhenhang, et autres
Publié: (2026)
par: Shang, Zhenhang, et autres
Publié: (2026)
How LoRA Remembers? A Parametric Memory Law for LLM Finetuning
par: Xu, Ziwen, et autres
Publié: (2026)
par: Xu, Ziwen, et autres
Publié: (2026)
Guarding the Guardrails: A Taxonomy-Driven Approach to Jailbreak Detection
par: Giarrusso, Francesco, et autres
Publié: (2025)
par: Giarrusso, Francesco, et autres
Publié: (2025)
WebGuard: Building a Generalizable Guardrail for Web Agents
par: Zheng, Boyuan, et autres
Publié: (2025)
par: Zheng, Boyuan, et autres
Publié: (2025)
BiasGuard: Guardrailing Fairness in Machine Learning Production Systems
par: Cohen-Inger, Nurit, et autres
Publié: (2025)
par: Cohen-Inger, Nurit, et autres
Publié: (2025)
ThinkGuard: Deliberative Slow Thinking Leads to Cautious Guardrails
par: Wen, Xiaofei, et autres
Publié: (2025)
par: Wen, Xiaofei, et autres
Publié: (2025)
AlignGuard: Scalable Safety Alignment for Text-to-Image Generation
par: Liu, Runtao, et autres
Publié: (2024)
par: Liu, Runtao, et autres
Publié: (2024)
A Causal Perspective for Enhancing Jailbreak Attack and Defense
par: Pan, Licheng, et autres
Publié: (2026)
par: Pan, Licheng, et autres
Publié: (2026)
TEMPO: Temporal Enforcement via Mode-Separated Policy Optimization for Trustworthy LLM Backtesting
par: Zhang, Zeyu, et autres
Publié: (2026)
par: Zhang, Zeyu, et autres
Publié: (2026)
Towards Real-world Debiasing: Rethinking Evaluation, Challenge, and Solution
par: Kuang, Peng, et autres
Publié: (2024)
par: Kuang, Peng, et autres
Publié: (2024)
CareGuardAI: Context-Aware Multi-Agent Guardrails for Clinical Safety & Hallucination Mitigation in Patient-Facing LLMs
par: Nasarian, Elham, et autres
Publié: (2026)
par: Nasarian, Elham, et autres
Publié: (2026)
ALMGuard: Safety Shortcuts and Where to Find Them as Guardrails for Audio-Language Models
par: Jin, Weifei, et autres
Publié: (2025)
par: Jin, Weifei, et autres
Publié: (2025)
MindGuard: Guardrail Classifiers for Multi-Turn Mental Health Support
par: Farinhas, António, et autres
Publié: (2026)
par: Farinhas, António, et autres
Publié: (2026)
SentGuard: Sentence-Level Streaming Guardrails for Large Language Models
par: Yu, Jiaqi, et autres
Publié: (2026)
par: Yu, Jiaqi, et autres
Publié: (2026)
Bag of Tricks for Subverting Reasoning-based Safety Guardrails
par: Chen, Shuo, et autres
Publié: (2025)
par: Chen, Shuo, et autres
Publié: (2025)
OceanGPT: A Large Language Model for Ocean Science Tasks
par: Bi, Zhen, et autres
Publié: (2023)
par: Bi, Zhen, et autres
Publié: (2023)
S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models
par: Yuan, Xiaohan, et autres
Publié: (2024)
par: Yuan, Xiaohan, et autres
Publié: (2024)
TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts
par: Chu, Hua-Rong, et autres
Publié: (2026)
par: Chu, Hua-Rong, et autres
Publié: (2026)
Perplexity-Aware Data Scaling Law: Perplexity Landscapes Predict Performance for Continual Pre-training
par: Liu, Lei, et autres
Publié: (2025)
par: Liu, Lei, et autres
Publié: (2025)
Documents similaires
-
Make LLM Learn to Synthesize from Streaming Experiences through Feedback
par: Hu, Zhenlin, et autres
Publié: (2026) -
Robust and Generalizable Safety Steering for Text-to-Image Diffusion Transformers
par: Xue, Zihao, et autres
Publié: (2026) -
OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning
par: Zhu, Boyu, et autres
Publié: (2025) -
A Causal Explainable Guardrails for Large Language Models
par: Chu, Zhixuan, et autres
Publié: (2024) -
Spatial Knowledge Graph-Guided Multimodal Synthesis
par: Xue, Yida, et autres
Publié: (2025)