OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhu, Boyu, Wen, Xiaofei, Mo, Wenjie Jacky, Zhu, Tinghui, Xie, Yanan, Qi, Peng, Chen, Muhao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ThinkGuard: Deliberative Slow Thinking Leads to Cautious Guardrails
por: Wen, Xiaofei, et al.
Publicado: (2025)
por: Wen, Xiaofei, et al.
Publicado: (2025)
Triaging Threats to Specialized Guardrails
por: Mo, Wenjie Jacky, et al.
Publicado: (2026)
por: Mo, Wenjie Jacky, et al.
Publicado: (2026)
GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio
por: Zhu, Zhenhao, et al.
Publicado: (2026)
por: Zhu, Zhenhao, et al.
Publicado: (2026)
Is Extending Modality The Right Path Towards Omni-Modality?
por: Zhu, Tinghui, et al.
Publicado: (2025)
por: Zhu, Tinghui, et al.
Publicado: (2025)
Rethinking Backdoor Detection Evaluation for Language Models
por: Yan, Jun, et al.
Publicado: (2024)
por: Yan, Jun, et al.
Publicado: (2024)
OpenGuardrails: A Configurable, Unified, and Scalable Guardrails Platform for Large Language Models
por: Wang, Thomas, et al.
Publicado: (2025)
por: Wang, Thomas, et al.
Publicado: (2025)
ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models
por: Zhao, Yunhan, et al.
Publicado: (2026)
por: Zhao, Yunhan, et al.
Publicado: (2026)
Omni-Safety under Cross-Modality Conflict: Vulnerabilities, Dynamics Mechanisms and Efficient Alignment
por: Wang, Kun, et al.
Publicado: (2026)
por: Wang, Kun, et al.
Publicado: (2026)
Bag of Tricks for Subverting Reasoning-based Safety Guardrails
por: Chen, Shuo, et al.
Publicado: (2025)
por: Chen, Shuo, et al.
Publicado: (2025)
InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models
por: Li, Hao, et al.
Publicado: (2024)
por: Li, Hao, et al.
Publicado: (2024)
RTD-Guard: A Black-Box Textual Adversarial Detection Framework via Replacement Token Detection
por: Zhu, He, et al.
Publicado: (2026)
por: Zhu, He, et al.
Publicado: (2026)
Black-Box Guardrail Reverse-engineering Attack
por: Yao, Hongwei, et al.
Publicado: (2025)
por: Yao, Hongwei, et al.
Publicado: (2025)
DualGuard: Dual-stream Large Language Model Watermarking Defense against Paraphrase and Spoofing Attack
por: Li, Hao, et al.
Publicado: (2025)
por: Li, Hao, et al.
Publicado: (2025)
Interpretable LLM Guardrails via Sparse Representation Steering
por: He, Zeqing, et al.
Publicado: (2025)
por: He, Zeqing, et al.
Publicado: (2025)
BraveGuard: From Open-World Threats to Safer Computer-Use Agents
por: Feng, Yunhao, et al.
Publicado: (2026)
por: Feng, Yunhao, et al.
Publicado: (2026)
TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts
por: Chu, Hua-Rong, et al.
Publicado: (2026)
por: Chu, Hua-Rong, et al.
Publicado: (2026)
GLiGuard: Schema-Conditioned Classification for LLM Safeguard
por: Zaratiana, Urchade, et al.
Publicado: (2026)
por: Zaratiana, Urchade, et al.
Publicado: (2026)
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
por: Shen, Guobin, et al.
Publicado: (2025)
por: Shen, Guobin, et al.
Publicado: (2025)
OneShield -- the Next Generation of LLM Guardrails
por: DeLuca, Chad, et al.
Publicado: (2025)
por: DeLuca, Chad, et al.
Publicado: (2025)
ConfGuard: A Simple and Effective Backdoor Detection for Large Language Models
por: Wang, Zihan, et al.
Publicado: (2025)
por: Wang, Zihan, et al.
Publicado: (2025)
PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails
por: Mangaokar, Neal, et al.
Publicado: (2024)
por: Mangaokar, Neal, et al.
Publicado: (2024)
Proof-of-Guardrail in AI Agents and What (Not) to Trust from It
por: Jin, Xisen, et al.
Publicado: (2026)
por: Jin, Xisen, et al.
Publicado: (2026)
NeuroFilter: Privacy Guardrails for Conversational LLM Agents
por: Das, Saswat, et al.
Publicado: (2026)
por: Das, Saswat, et al.
Publicado: (2026)
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
por: Wang, Jiongxiao, et al.
Publicado: (2024)
por: Wang, Jiongxiao, et al.
Publicado: (2024)
SGuard-v1: Safety Guardrail for Large Language Models
por: Lee, JoonHo, et al.
Publicado: (2025)
por: Lee, JoonHo, et al.
Publicado: (2025)
Omni-IML: Towards Unified Image Manipulation Localization
por: Qu, Chenfan, et al.
Publicado: (2024)
por: Qu, Chenfan, et al.
Publicado: (2024)
Auto-Tuning Safety Guardrails for Black-Box Large Language Models
por: Abdulkadir, Perry
Publicado: (2025)
por: Abdulkadir, Perry
Publicado: (2025)
FATH: Authentication-based Test-time Defense against Indirect Prompt Injection Attacks
por: Wang, Jiongxiao, et al.
Publicado: (2024)
por: Wang, Jiongxiao, et al.
Publicado: (2024)
OmniSafeBench-MM: A Unified Benchmark and Toolbox for Multimodal Jailbreak Attack-Defense Evaluation
por: Jia, Xiaojun, et al.
Publicado: (2025)
por: Jia, Xiaojun, et al.
Publicado: (2025)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
por: Tong, Terry, et al.
Publicado: (2025)
por: Tong, Terry, et al.
Publicado: (2025)
ADVERSA: Measuring Multi-Turn Guardrail Degradation and Judge Reliability in Large Language Models
por: Owiredu-Ashley, Harry
Publicado: (2026)
por: Owiredu-Ashley, Harry
Publicado: (2026)
Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
por: Wei, Zeming, et al.
Publicado: (2023)
por: Wei, Zeming, et al.
Publicado: (2023)
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures
por: Zhou, Yukai, et al.
Publicado: (2025)
por: Zhou, Yukai, et al.
Publicado: (2025)
From Trade-off to Synergy: A Versatile Symbiotic Watermarking Framework for Large Language Models
por: Wang, Yidan, et al.
Publicado: (2025)
por: Wang, Yidan, et al.
Publicado: (2025)
WorldCup Sampling for Multi-bit LLM Watermarking
por: Wang, Yidan, et al.
Publicado: (2026)
por: Wang, Yidan, et al.
Publicado: (2026)
LLMGuard: Guarding Against Unsafe LLM Behavior
por: Goyal, Shubh, et al.
Publicado: (2024)
por: Goyal, Shubh, et al.
Publicado: (2024)
Mitigating Backdoor Threats to Large Language Models: Advancement and Challenges
por: Liu, Qin, et al.
Publicado: (2024)
por: Liu, Qin, et al.
Publicado: (2024)
Learning Efficient Guardrails for Compliance
por: Wen, Xiaofei, et al.
Publicado: (2025)
por: Wen, Xiaofei, et al.
Publicado: (2025)
From Retrieval to Reasoning: A Framework for Cyber Threat Intelligence NER with Explicit and Adaptive Instructions
por: Peng, Jiaren, et al.
Publicado: (2025)
por: Peng, Jiaren, et al.
Publicado: (2025)
Fact2Fiction: Targeted Poisoning Attack to Agentic Fact-checking System
por: He, Haorui, et al.
Publicado: (2025)
por: He, Haorui, et al.
Publicado: (2025)
Ejemplares similares
-
ThinkGuard: Deliberative Slow Thinking Leads to Cautious Guardrails
por: Wen, Xiaofei, et al.
Publicado: (2025) -
Triaging Threats to Specialized Guardrails
por: Mo, Wenjie Jacky, et al.
Publicado: (2026) -
GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio
por: Zhu, Zhenhao, et al.
Publicado: (2026) -
Is Extending Modality The Right Path Towards Omni-Modality?
por: Zhu, Tinghui, et al.
Publicado: (2025) -
Rethinking Backdoor Detection Evaluation for Language Models
por: Yan, Jun, et al.
Publicado: (2024)