X-Guard: Multilingual Guard Agent for Content Moderation
Fuente:
arXiv
Salvato in:
| Autori principali: | Upadhayay, Bibek, Behzadan, Vahid, D, Ph. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Sandwich attack: Multi-language Mixture Adaptive Attack on LLMs
di: Upadhayay, Bibek, et al.
Pubblicazione: (2024)
di: Upadhayay, Bibek, et al.
Pubblicazione: (2024)
TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in LLMs through Translation-Assisted Chain-of-Thought Processes
di: Upadhayay, Bibek, et al.
Pubblicazione: (2023)
di: Upadhayay, Bibek, et al.
Pubblicazione: (2023)
AIRGuard: Guarding Agent Actions with Runtime Authority Control
di: Qin, Suliu, et al.
Pubblicazione: (2026)
di: Qin, Suliu, et al.
Pubblicazione: (2026)
Sentra-Guard: A Real-Time Multilingual Defense Against Adversarial LLM Prompts
di: Hasan, Md. Mehedi, et al.
Pubblicazione: (2025)
di: Hasan, Md. Mehedi, et al.
Pubblicazione: (2025)
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
di: Yuan, Lingzhi, et al.
Pubblicazione: (2025)
di: Yuan, Lingzhi, et al.
Pubblicazione: (2025)
AgentGuard: Repurposing Agentic Orchestrator for Safety Evaluation of Tool Orchestration
di: Chen, Jizhou, et al.
Pubblicazione: (2025)
di: Chen, Jizhou, et al.
Pubblicazione: (2025)
RouteGuard: Internal-Signal Detection of Skill Poisoning in LLM Agents
di: Xiao, Wenjie, et al.
Pubblicazione: (2026)
di: Xiao, Wenjie, et al.
Pubblicazione: (2026)
RL-Based Method for Benchmarking the Adversarial Resilience and Robustness of Deep Reinforcement Learning Policies
di: Behzadan, Vahid, et al.
Pubblicazione: (2019)
di: Behzadan, Vahid, et al.
Pubblicazione: (2019)
SnapGuard: Lightweight Prompt Injection Detection for Screenshot-Based Web Agents
di: Du, Mengyao, et al.
Pubblicazione: (2026)
di: Du, Mengyao, et al.
Pubblicazione: (2026)
CoT-Guard: Small Models for Strong Monitoring
di: Diwan, Nirav, et al.
Pubblicazione: (2026)
di: Diwan, Nirav, et al.
Pubblicazione: (2026)
A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory
di: Wei, Qianshan, et al.
Pubblicazione: (2025)
di: Wei, Qianshan, et al.
Pubblicazione: (2025)
MultiPhishGuard: An Explainable and Adaptive Multi-Agent LLM System for Phishing Email Detection
di: Xue, Yinuo, et al.
Pubblicazione: (2025)
di: Xue, Yinuo, et al.
Pubblicazione: (2025)
CoopGuard: Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Round Attacks
di: Li, Siyuan, et al.
Pubblicazione: (2026)
di: Li, Siyuan, et al.
Pubblicazione: (2026)
GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
di: Liu, Yue, et al.
Pubblicazione: (2025)
di: Liu, Yue, et al.
Pubblicazione: (2025)
CourtGuard: A Local, Multiagent Prompt Injection Classifier
di: Wu, Isaac, et al.
Pubblicazione: (2025)
di: Wu, Isaac, et al.
Pubblicazione: (2025)
ClawGuard: A Runtime Security Framework for Tool-Augmented LLM Agents Against Indirect Prompt Injection
di: Zhao, Wei, et al.
Pubblicazione: (2026)
di: Zhao, Wei, et al.
Pubblicazione: (2026)
MergeGuard: Efficient Thwarting of Trojan Attacks in Machine Learning Models
di: Shabgahi, Soheil Zibakhsh, et al.
Pubblicazione: (2025)
di: Shabgahi, Soheil Zibakhsh, et al.
Pubblicazione: (2025)
VibeGuard: A Security Gate Framework for AI-Generated Code
di: Xie, Ying
Pubblicazione: (2026)
di: Xie, Ying
Pubblicazione: (2026)
Grimlock: Guarding High-Agency Systems with eBPF and Attested Channels
di: Wu, Qiancheng, et al.
Pubblicazione: (2026)
di: Wu, Qiancheng, et al.
Pubblicazione: (2026)
Privacy Guard & Token Parsimony by Prompt and Context Handling and LLM Routing
di: Langiu, Alessio
Pubblicazione: (2026)
di: Langiu, Alessio
Pubblicazione: (2026)
QGuard:Question-based Zero-shot Guard for Multi-modal LLM Safety
di: Lee, Taegyeong, et al.
Pubblicazione: (2025)
di: Lee, Taegyeong, et al.
Pubblicazione: (2025)
CrossGuard: Safeguarding MLLMs against Joint-Modal Implicit Malicious Attacks
di: Zhang, Xu, et al.
Pubblicazione: (2025)
di: Zhang, Xu, et al.
Pubblicazione: (2025)
TraceGuard: Structured Multi-Dimensional Monitoring as a Collusion-Resistant Control Protocol
di: Nguyen, Khanh Linh, et al.
Pubblicazione: (2026)
di: Nguyen, Khanh Linh, et al.
Pubblicazione: (2026)
TransLinkGuard: Safeguarding Transformer Models Against Model Stealing in Edge Deployment
di: Li, Qinfeng, et al.
Pubblicazione: (2024)
di: Li, Qinfeng, et al.
Pubblicazione: (2024)
TrajGuard: Streaming Hidden-state Trajectory Detection for Decoding-time Jailbreak Defense
di: Liu, Cheng, et al.
Pubblicazione: (2026)
di: Liu, Cheng, et al.
Pubblicazione: (2026)
ExplainableGuard: Interpretable Adversarial Defense for Large Language Models Using Chain-of-Thought Reasoning
di: Guan, Shaowei, et al.
Pubblicazione: (2025)
di: Guan, Shaowei, et al.
Pubblicazione: (2025)
Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection
di: Lin, Lixing, et al.
Pubblicazione: (2026)
di: Lin, Lixing, et al.
Pubblicazione: (2026)
GuardReasoner: Towards Reasoning-based LLM Safeguards
di: Liu, Yue, et al.
Pubblicazione: (2025)
di: Liu, Yue, et al.
Pubblicazione: (2025)
DistillGuard: Evaluating Defenses Against LLM Knowledge Distillation
di: Jiang, Bo
Pubblicazione: (2026)
di: Jiang, Bo
Pubblicazione: (2026)
MCP-Guard: A Multi-Stage Defense-in-Depth Framework for Securing Model Context Protocol in Agentic AI
di: Xing, Wenpeng, et al.
Pubblicazione: (2025)
di: Xing, Wenpeng, et al.
Pubblicazione: (2025)
TinyGuard:A lightweight Byzantine Defense for Resource-Constrained Federated Learning via Statistical Update Fingerprints
di: Mahdavi, Ali, et al.
Pubblicazione: (2026)
di: Mahdavi, Ali, et al.
Pubblicazione: (2026)
AdaptiveGuard: Towards Adaptive Runtime Safety for LLM-Powered Software
di: Yang, Rui, et al.
Pubblicazione: (2025)
di: Yang, Rui, et al.
Pubblicazione: (2025)
AEGIS : Automated Co-Evolutionary Framework for Guarding Prompt Injections Schema
di: Liu, Ting-Chun, et al.
Pubblicazione: (2025)
di: Liu, Ting-Chun, et al.
Pubblicazione: (2025)
Adversarial Distilled Retrieval-Augmented Guarding Model for Online Malicious Intent Detection
di: Guo, Yihao, et al.
Pubblicazione: (2025)
di: Guo, Yihao, et al.
Pubblicazione: (2025)
DeepGuard: Secure Code Generation via Multi-Layer Semantic Aggregation
di: Huang, Li, et al.
Pubblicazione: (2026)
di: Huang, Li, et al.
Pubblicazione: (2026)
PropGuard: Safeguarding LLM-MAS via Propagation-Aware Exploration and Remediation
di: Yan, Bingyu, et al.
Pubblicazione: (2026)
di: Yan, Bingyu, et al.
Pubblicazione: (2026)
When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models
di: Hossain, Ismail, et al.
Pubblicazione: (2026)
di: Hossain, Ismail, et al.
Pubblicazione: (2026)
InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models
di: Li, Hao, et al.
Pubblicazione: (2024)
di: Li, Hao, et al.
Pubblicazione: (2024)
Guarding Your Conversations: Privacy Gatekeepers for Secure Interactions with Cloud-Based AI Models
di: Uzor, GodsGift, et al.
Pubblicazione: (2025)
di: Uzor, GodsGift, et al.
Pubblicazione: (2025)
MoJE: Mixture of Jailbreak Experts, Naive Tabular Classifiers as Guard for Prompt Attacks
di: Cornacchia, Giandomenico, et al.
Pubblicazione: (2024)
di: Cornacchia, Giandomenico, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Sandwich attack: Multi-language Mixture Adaptive Attack on LLMs
di: Upadhayay, Bibek, et al.
Pubblicazione: (2024) -
TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in LLMs through Translation-Assisted Chain-of-Thought Processes
di: Upadhayay, Bibek, et al.
Pubblicazione: (2023) -
AIRGuard: Guarding Agent Actions with Runtime Authority Control
di: Qin, Suliu, et al.
Pubblicazione: (2026) -
Sentra-Guard: A Real-Time Multilingual Defense Against Adversarial LLM Prompts
di: Hasan, Md. Mehedi, et al.
Pubblicazione: (2025) -
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
di: Yuan, Lingzhi, et al.
Pubblicazione: (2025)