Salvato in:
| Autori principali: | Wang, Kongxin, Zhang, Jie, Qi, Peigui, Tang, Kunsheng, Zhang, Tianwei, Zhou, Wenbo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2508.02476 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SafeGuider: Robust and Practical Content Safety Control for Text-to-Image Models
di: Qi, Peigui, et al.
Pubblicazione: (2025)
di: Qi, Peigui, et al.
Pubblicazione: (2025)
Poly-Guard: Massive Multi-Domain Safety Policy-Grounded Guardrail Dataset
di: Kang, Mintong, et al.
Pubblicazione: (2025)
di: Kang, Mintong, et al.
Pubblicazione: (2025)
State-Dependent Safety Failures in Multi-Turn Language Model Interaction
di: Li, Pengcheng, et al.
Pubblicazione: (2026)
di: Li, Pengcheng, et al.
Pubblicazione: (2026)
Invisibility Cloak: Disappearance under Human Pose Estimation via Backdoor Attacks
di: Zhang, Minxing, et al.
Pubblicazione: (2024)
di: Zhang, Minxing, et al.
Pubblicazione: (2024)
ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models
di: Zhao, Yunhan, et al.
Pubblicazione: (2026)
di: Zhao, Yunhan, et al.
Pubblicazione: (2026)
Turning Your Strength into Watermark: Watermarking Large Language Model via Knowledge Injection
di: Li, Shuai, et al.
Pubblicazione: (2023)
di: Li, Shuai, et al.
Pubblicazione: (2023)
Character as a Latent Variable in Large Language Models: A Mechanistic Account of Emergent Misalignment and Conditional Safety Failures
di: Su, Yanghao, et al.
Pubblicazione: (2026)
di: Su, Yanghao, et al.
Pubblicazione: (2026)
GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio
di: Zhu, Zhenhao, et al.
Pubblicazione: (2026)
di: Zhu, Zhenhao, et al.
Pubblicazione: (2026)
CipherGuard: Compiler-aided Mitigation against Ciphertext Side-channel Attacks
di: Jiang, Ke, et al.
Pubblicazione: (2025)
di: Jiang, Ke, et al.
Pubblicazione: (2025)
ThinkGuard: Deliberative Slow Thinking Leads to Cautious Guardrails
di: Wen, Xiaofei, et al.
Pubblicazione: (2025)
di: Wen, Xiaofei, et al.
Pubblicazione: (2025)
BURN: Backdoor Unlearning via Adversarial Boundary Analysis
di: Su, Yanghao, et al.
Pubblicazione: (2025)
di: Su, Yanghao, et al.
Pubblicazione: (2025)
AquaLoRA: Toward White-box Protection for Customized Stable Diffusion Models via Watermark LoRA
di: Feng, Weitao, et al.
Pubblicazione: (2024)
di: Feng, Weitao, et al.
Pubblicazione: (2024)
On the Account Security Risks Posed by Password Strength Meters
di: Xu, Ming, et al.
Pubblicazione: (2025)
di: Xu, Ming, et al.
Pubblicazione: (2025)
From AI-Generated Content to Agentic Action: Security and Safety Threats in Generative AI
di: Zhang, Zelin, et al.
Pubblicazione: (2026)
di: Zhang, Zelin, et al.
Pubblicazione: (2026)
ALMGuard: Safety Shortcuts and Where to Find Them as Guardrails for Audio-Language Models
di: Jin, Weifei, et al.
Pubblicazione: (2025)
di: Jin, Weifei, et al.
Pubblicazione: (2025)
SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety
di: Liu, Zhe, et al.
Pubblicazione: (2026)
di: Liu, Zhe, et al.
Pubblicazione: (2026)
Bag of Tricks for Subverting Reasoning-based Safety Guardrails
di: Chen, Shuo, et al.
Pubblicazione: (2025)
di: Chen, Shuo, et al.
Pubblicazione: (2025)
InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models
di: Li, Hao, et al.
Pubblicazione: (2024)
di: Li, Hao, et al.
Pubblicazione: (2024)
Siren Song: Manipulating Pose Estimation in XR Headsets Using Acoustic Attacks
di: Huang, Zijian, et al.
Pubblicazione: (2025)
di: Huang, Zijian, et al.
Pubblicazione: (2025)
OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning
di: Zhu, Boyu, et al.
Pubblicazione: (2025)
di: Zhu, Boyu, et al.
Pubblicazione: (2025)
OneShield -- the Next Generation of LLM Guardrails
di: DeLuca, Chad, et al.
Pubblicazione: (2025)
di: DeLuca, Chad, et al.
Pubblicazione: (2025)
A Comparative Evaluation of AI Agent Security Guardrails
di: Li, Qi, et al.
Pubblicazione: (2026)
di: Li, Qi, et al.
Pubblicazione: (2026)
InferDPT: Privacy-Preserving Inference for Closed-box Large Language Model
di: Tong, Meng, et al.
Pubblicazione: (2023)
di: Tong, Meng, et al.
Pubblicazione: (2023)
TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts
di: Chu, Hua-Rong, et al.
Pubblicazione: (2026)
di: Chu, Hua-Rong, et al.
Pubblicazione: (2026)
ConfGuard: A Simple and Effective Backdoor Detection for Large Language Models
di: Wang, Zihan, et al.
Pubblicazione: (2025)
di: Wang, Zihan, et al.
Pubblicazione: (2025)
Robust-Wide: Robust Watermarking against Instruction-driven Image Editing
di: Hu, Runyi, et al.
Pubblicazione: (2024)
di: Hu, Runyi, et al.
Pubblicazione: (2024)
Provably Secure Agent Guardrail
di: Wu, Benlong, et al.
Pubblicazione: (2026)
di: Wu, Benlong, et al.
Pubblicazione: (2026)
TraceGuard: Process-Guided Firewall against Reasoning Backdoors in Large Language Models
di: Guo, Zhen, et al.
Pubblicazione: (2026)
di: Guo, Zhen, et al.
Pubblicazione: (2026)
Peering Behind the Shield: Guardrail Identification in Large Language Models
di: Yang, Ziqing, et al.
Pubblicazione: (2025)
di: Yang, Ziqing, et al.
Pubblicazione: (2025)
Investigating Threats Posed by SMS Origin Spoofing to IoT Devices
di: Tsunoda, Akaki
Pubblicazione: (2023)
di: Tsunoda, Akaki
Pubblicazione: (2023)
The Gradient Puppeteer: Adversarial Domination in Gradient Leakage Attacks through Model Poisoning
di: Xiang, Kunlan, et al.
Pubblicazione: (2025)
di: Xiang, Kunlan, et al.
Pubblicazione: (2025)
JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks
di: Zhang, Xiaoyu, et al.
Pubblicazione: (2023)
di: Zhang, Xiaoyu, et al.
Pubblicazione: (2023)
OpenGuardrails: A Configurable, Unified, and Scalable Guardrails Platform for Large Language Models
di: Wang, Thomas, et al.
Pubblicazione: (2025)
di: Wang, Thomas, et al.
Pubblicazione: (2025)
Hoist with His Own Petard: Inducing Guardrails to Facilitate Denial-of-Service Attacks on Retrieval-Augmented Generation of LLMs
di: Suo, Pan, et al.
Pubblicazione: (2025)
di: Suo, Pan, et al.
Pubblicazione: (2025)
PSRT: Accelerating LRM-based Guard Models via Prefilled Safe Reasoning Traces
di: Zhao, Jiawei, et al.
Pubblicazione: (2025)
di: Zhao, Jiawei, et al.
Pubblicazione: (2025)
GLiNER Guard: Unified Encoder Family for Production LLM Safety and Privacy
di: Minko, Bogdan, et al.
Pubblicazione: (2026)
di: Minko, Bogdan, et al.
Pubblicazione: (2026)
Oedipus: LLM-enchanced Reasoning CAPTCHA Solver
di: Deng, Gelei, et al.
Pubblicazione: (2024)
di: Deng, Gelei, et al.
Pubblicazione: (2024)
Interpretable LLM Guardrails via Sparse Representation Steering
di: He, Zeqing, et al.
Pubblicazione: (2025)
di: He, Zeqing, et al.
Pubblicazione: (2025)
Pandora: Jailbreak GPTs by Retrieval Augmented Generation Poisoning
di: Deng, Gelei, et al.
Pubblicazione: (2024)
di: Deng, Gelei, et al.
Pubblicazione: (2024)
SSD: A State-based Stealthy Backdoor Attack For Navigation System in UAV Route Planning
di: Wang, Zhaoxuan, et al.
Pubblicazione: (2025)
di: Wang, Zhaoxuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
SafeGuider: Robust and Practical Content Safety Control for Text-to-Image Models
di: Qi, Peigui, et al.
Pubblicazione: (2025) -
Poly-Guard: Massive Multi-Domain Safety Policy-Grounded Guardrail Dataset
di: Kang, Mintong, et al.
Pubblicazione: (2025) -
State-Dependent Safety Failures in Multi-Turn Language Model Interaction
di: Li, Pengcheng, et al.
Pubblicazione: (2026) -
Invisibility Cloak: Disappearance under Human Pose Estimation via Backdoor Attacks
di: Zhang, Minxing, et al.
Pubblicazione: (2024) -
ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models
di: Zhao, Yunhan, et al.
Pubblicazione: (2026)