$R^2$-Guard: Robust Reasoning Enabled LLM Guardrail via Knowledge-Enhanced Logical Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | Kang, Mintong, Li, Bo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
COLEP: Certifiably Robust Learning-Reasoning Conformal Prediction via Probabilistic Circuits
di: Kang, Mintong, et al.
Pubblicazione: (2024)
di: Kang, Mintong, et al.
Pubblicazione: (2024)
AudioGuard: Toward Comprehensive Audio Safety Protection Across Diverse Threat Models
di: Kang, Mintong, et al.
Pubblicazione: (2026)
di: Kang, Mintong, et al.
Pubblicazione: (2026)
Logical Structure as Knowledge: Enhancing LLM Reasoning via Structured Logical Knowledge Density Estimation
di: Bi, Zhen, et al.
Pubblicazione: (2025)
di: Bi, Zhen, et al.
Pubblicazione: (2025)
CodeGuard: Improving LLM Guardrails in CS Education
di: Raihan, Nishat, et al.
Pubblicazione: (2026)
di: Raihan, Nishat, et al.
Pubblicazione: (2026)
OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning
di: Zhu, Boyu, et al.
Pubblicazione: (2025)
di: Zhu, Boyu, et al.
Pubblicazione: (2025)
KnowGraph: Knowledge-Enabled Anomaly Detection via Logical Reasoning on Graph Data
di: Zhou, Andy, et al.
Pubblicazione: (2024)
di: Zhou, Andy, et al.
Pubblicazione: (2024)
GuardReasoner: Towards Reasoning-based LLM Safeguards
di: Liu, Yue, et al.
Pubblicazione: (2025)
di: Liu, Yue, et al.
Pubblicazione: (2025)
DiffAttack: Evasion Attacks Against Diffusion-Based Adversarial Purification
di: Kang, Mintong, et al.
Pubblicazione: (2023)
di: Kang, Mintong, et al.
Pubblicazione: (2023)
Certifiably Byzantine-Robust Federated Conformal Prediction
di: Kang, Mintong, et al.
Pubblicazione: (2024)
di: Kang, Mintong, et al.
Pubblicazione: (2024)
Safety Guardrails for LLM-Enabled Robots
di: Ravichandran, Zachary, et al.
Pubblicazione: (2025)
di: Ravichandran, Zachary, et al.
Pubblicazione: (2025)
GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
di: Liu, Yue, et al.
Pubblicazione: (2025)
di: Liu, Yue, et al.
Pubblicazione: (2025)
Robust Answers, Fragile Logic: Probing the Decoupling Hypothesis in LLM Reasoning
di: Jiang, Enyi, et al.
Pubblicazione: (2025)
di: Jiang, Enyi, et al.
Pubblicazione: (2025)
Safety Through Reasoning: An Empirical Study of Reasoning Guardrail Models
di: Sreedhar, Makesh Narsimhan, et al.
Pubblicazione: (2025)
di: Sreedhar, Makesh Narsimhan, et al.
Pubblicazione: (2025)
AdvWave: Stealthy Adversarial Jailbreak Attack against Large Audio-Language Models
di: Kang, Mintong, et al.
Pubblicazione: (2024)
di: Kang, Mintong, et al.
Pubblicazione: (2024)
Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection
di: Lin, Lixing, et al.
Pubblicazione: (2026)
di: Lin, Lixing, et al.
Pubblicazione: (2026)
Reliable Reasoning Path: Distilling Effective Guidance for LLM Reasoning with Knowledge Graphs
di: Xiao, Yilin, et al.
Pubblicazione: (2025)
di: Xiao, Yilin, et al.
Pubblicazione: (2025)
Logics-STEM: Empowering LLM Reasoning via Failure-Driven Post-Training and Document Knowledge Enhancement
di: Xu, Mingyu, et al.
Pubblicazione: (2026)
di: Xu, Mingyu, et al.
Pubblicazione: (2026)
DRAGON: Guard LLM Unlearning in Context via Negative Detection and Reasoning
di: Wang, Yaxuan, et al.
Pubblicazione: (2025)
di: Wang, Yaxuan, et al.
Pubblicazione: (2025)
MindGuard: Guardrail Classifiers for Multi-Turn Mental Health Support
di: Farinhas, António, et al.
Pubblicazione: (2026)
di: Farinhas, António, et al.
Pubblicazione: (2026)
LPG: Balancing Efficiency and Policy Reasoning in Latent Policy Guardrails
di: Li, Nanxi, et al.
Pubblicazione: (2026)
di: Li, Nanxi, et al.
Pubblicazione: (2026)
ChemDFM-R: A Chemical Reasoning LLM Enhanced with Atomized Chemical Knowledge
di: Zhao, Zihan, et al.
Pubblicazione: (2025)
di: Zhao, Zihan, et al.
Pubblicazione: (2025)
DistillGuard: Evaluating Defenses Against LLM Knowledge Distillation
di: Jiang, Bo
Pubblicazione: (2026)
di: Jiang, Bo
Pubblicazione: (2026)
An Enhanced Prompt-Based LLM Reasoning Scheme via Knowledge Graph-Integrated Collaboration
di: Li, Yihao, et al.
Pubblicazione: (2024)
di: Li, Yihao, et al.
Pubblicazione: (2024)
Controllable Logical Hypothesis Generation for Abductive Reasoning in Knowledge Graphs
di: Gao, Yisen, et al.
Pubblicazione: (2025)
di: Gao, Yisen, et al.
Pubblicazione: (2025)
Guarding the Guardrails: A Taxonomy-Driven Approach to Jailbreak Detection
di: Giarrusso, Francesco, et al.
Pubblicazione: (2025)
di: Giarrusso, Francesco, et al.
Pubblicazione: (2025)
Evaluating Robustness of Reasoning Models on Parameterized Logical Problems
di: Es-sebbani, Naïm, et al.
Pubblicazione: (2026)
di: Es-sebbani, Naïm, et al.
Pubblicazione: (2026)
Reason from Fallacy: Enhancing Large Language Models' Logical Reasoning through Logical Fallacy Understanding
di: Li, Yanda, et al.
Pubblicazione: (2024)
di: Li, Yanda, et al.
Pubblicazione: (2024)
ChatRule: Mining Logical Rules with Large Language Models for Knowledge Graph Reasoning
di: Luo, Linhao, et al.
Pubblicazione: (2023)
di: Luo, Linhao, et al.
Pubblicazione: (2023)
Assessing LLM Reasoning Steps via Principal Knowledge Grounding
di: Hwang, Hyeon, et al.
Pubblicazione: (2025)
di: Hwang, Hyeon, et al.
Pubblicazione: (2025)
Logical Reasoning with Relation Network for Inductive Knowledge Graph Completion
di: Zhang, Qinggang, et al.
Pubblicazione: (2024)
di: Zhang, Qinggang, et al.
Pubblicazione: (2024)
RAG Makes Guardrails Unsafe? Investigating Robustness of Guardrails under RAG-style Contexts
di: She, Yining, et al.
Pubblicazione: (2025)
di: She, Yining, et al.
Pubblicazione: (2025)
Self-Guard: Defending Large Reasoning Models via enhanced self-reflection
di: Zheng, Jingnan, et al.
Pubblicazione: (2026)
di: Zheng, Jingnan, et al.
Pubblicazione: (2026)
InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models
di: Li, Hao, et al.
Pubblicazione: (2024)
di: Li, Hao, et al.
Pubblicazione: (2024)
BiasGuard: Guardrailing Fairness in Machine Learning Production Systems
di: Cohen-Inger, Nurit, et al.
Pubblicazione: (2025)
di: Cohen-Inger, Nurit, et al.
Pubblicazione: (2025)
VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection
di: Nie, Yuzhou, et al.
Pubblicazione: (2025)
di: Nie, Yuzhou, et al.
Pubblicazione: (2025)
Traceable Knowledge Graph Reasoning Enables LLM-Assisted Decision Support for Industrial VOCs in the Steel Industry
di: Su, Changqing, et al.
Pubblicazione: (2026)
di: Su, Changqing, et al.
Pubblicazione: (2026)
Adaptive Selection of Symbolic Languages for Improving LLM Logical Reasoning
di: Wang, Xiangyu, et al.
Pubblicazione: (2025)
di: Wang, Xiangyu, et al.
Pubblicazione: (2025)
Towards Unifying Perceptual Reasoning and Logical Reasoning
di: Kido, Hiroyuki
Pubblicazione: (2022)
di: Kido, Hiroyuki
Pubblicazione: (2022)
MemR$^3$: Memory Retrieval via Reflective Reasoning for LLM Agents
di: Du, Xingbo, et al.
Pubblicazione: (2025)
di: Du, Xingbo, et al.
Pubblicazione: (2025)
Demystifying the Roles of LLM Layers in Retrieval, Knowledge, and Reasoning
di: Song, Xinyuan, et al.
Pubblicazione: (2025)
di: Song, Xinyuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
COLEP: Certifiably Robust Learning-Reasoning Conformal Prediction via Probabilistic Circuits
di: Kang, Mintong, et al.
Pubblicazione: (2024) -
AudioGuard: Toward Comprehensive Audio Safety Protection Across Diverse Threat Models
di: Kang, Mintong, et al.
Pubblicazione: (2026) -
Logical Structure as Knowledge: Enhancing LLM Reasoning via Structured Logical Knowledge Density Estimation
di: Bi, Zhen, et al.
Pubblicazione: (2025) -
CodeGuard: Improving LLM Guardrails in CS Education
di: Raihan, Nishat, et al.
Pubblicazione: (2026) -
OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning
di: Zhu, Boyu, et al.
Pubblicazione: (2025)