Saved in:
| Main Authors: | Zaratiana, Urchade, Newhauser, Mary, Hurn-Maloney, George, Lewis, Ash |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.07982 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GLiNER2-PII: A Multilingual Model for Personally Identifiable Information Extraction
by: Zaratiana, Urchade, et al.
Published: (2026)
by: Zaratiana, Urchade, et al.
Published: (2026)
GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface
by: Zaratiana, Urchade, et al.
Published: (2025)
by: Zaratiana, Urchade, et al.
Published: (2025)
Pioneer Agent: Continual Improvement of Small Language Models in Production
by: Atreja, Dhruv, et al.
Published: (2026)
by: Atreja, Dhruv, et al.
Published: (2026)
GLiNER Guard: Unified Encoder Family for Production LLM Safety and Privacy
by: Minko, Bogdan, et al.
Published: (2026)
by: Minko, Bogdan, et al.
Published: (2026)
MaskSQL: Safeguarding Privacy for LLM-Based Text-to-SQL via Abstraction
by: Abedini, Sepideh, et al.
Published: (2025)
by: Abedini, Sepideh, et al.
Published: (2025)
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
by: Shen, Guobin, et al.
Published: (2025)
by: Shen, Guobin, et al.
Published: (2025)
LLMGuard: Guarding Against Unsafe LLM Behavior
by: Goyal, Shubh, et al.
Published: (2024)
by: Goyal, Shubh, et al.
Published: (2024)
TrinityGuard: A Unified Framework for Safeguarding Multi-Agent Systems
by: Wang, Kai, et al.
Published: (2026)
by: Wang, Kai, et al.
Published: (2026)
DistillGuard: Evaluating Defenses Against LLM Knowledge Distillation
by: Jiang, Bo
Published: (2026)
by: Jiang, Bo
Published: (2026)
On Calibration of LLM-based Guard Models for Reliable Content Moderation
by: Liu, Hongfu, et al.
Published: (2024)
by: Liu, Hongfu, et al.
Published: (2024)
BELLS: A Framework Towards Future Proof Benchmarks for the Evaluation of LLM Safeguards
by: Dorn, Diego, et al.
Published: (2024)
by: Dorn, Diego, et al.
Published: (2024)
MAGE: Safeguarding LLM Agents against Long-Horizon Threats via Shadow Memory
by: Wang, Yuhui, et al.
Published: (2026)
by: Wang, Yuhui, et al.
Published: (2026)
Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
by: Halawi, Danny, et al.
Published: (2024)
by: Halawi, Danny, et al.
Published: (2024)
Can Federated Learning Safeguard Private Data in LLM Training? Vulnerabilities, Attacks, and Defense Evaluation
by: Guo, Wenkai, et al.
Published: (2025)
by: Guo, Wenkai, et al.
Published: (2025)
CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-Tuning
by: Yi, Biao, et al.
Published: (2025)
by: Yi, Biao, et al.
Published: (2025)
BraveGuard: From Open-World Threats to Safer Computer-Use Agents
by: Feng, Yunhao, et al.
Published: (2026)
by: Feng, Yunhao, et al.
Published: (2026)
ConfGuard: A Simple and Effective Backdoor Detection for Large Language Models
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails
by: Mangaokar, Neal, et al.
Published: (2024)
by: Mangaokar, Neal, et al.
Published: (2024)
ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models
by: Zhao, Yunhan, et al.
Published: (2026)
by: Zhao, Yunhan, et al.
Published: (2026)
RTD-Guard: A Black-Box Textual Adversarial Detection Framework via Replacement Token Detection
by: Zhu, He, et al.
Published: (2026)
by: Zhu, He, et al.
Published: (2026)
DualGuard: Dual-stream Large Language Model Watermarking Defense against Paraphrase and Spoofing Attack
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
Protecting Users From Themselves: Safeguarding Contextual Privacy in Interactions with Conversational Agents
by: Ngong, Ivoline, et al.
Published: (2025)
by: Ngong, Ivoline, et al.
Published: (2025)
GuardReasoner: Towards Reasoning-based LLM Safeguards
by: Liu, Yue, et al.
Published: (2025)
by: Liu, Yue, et al.
Published: (2025)
LLM-based Privacy Data Augmentation Guided by Knowledge Distillation with a Distribution Tutor for Medical Text Classification
by: Song, Yiping, et al.
Published: (2024)
by: Song, Yiping, et al.
Published: (2024)
RED QUEEN: Safeguarding Large Language Models against Concealed Multi-Turn Jailbreaking
by: Jiang, Yifan, et al.
Published: (2024)
by: Jiang, Yifan, et al.
Published: (2024)
CVE-LLM : Ontology-Assisted Automatic Vulnerability Evaluation Using Large Language Models
by: Ghosh, Rikhiya, et al.
Published: (2025)
by: Ghosh, Rikhiya, et al.
Published: (2025)
InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Guarding Your Conversations: Privacy Gatekeepers for Secure Interactions with Cloud-Based AI Models
by: Uzor, GodsGift, et al.
Published: (2025)
by: Uzor, GodsGift, et al.
Published: (2025)
Towards Reliable and Practical LLM Security Evaluations via Bayesian Modelling
by: Llewellyn, Mary, et al.
Published: (2025)
by: Llewellyn, Mary, et al.
Published: (2025)
Why Safeguarded Ships Run Aground? Aligned Large Language Models' Safety Mechanisms Tend to Be Anchored in The Template Region
by: Leong, Chak Tou, et al.
Published: (2025)
by: Leong, Chak Tou, et al.
Published: (2025)
LLM Reinforcement in Context
by: Rivasseau, Thomas
Published: (2025)
by: Rivasseau, Thomas
Published: (2025)
Watermarking LLM Agent Trajectories
by: Meng, Wenlong, et al.
Published: (2026)
by: Meng, Wenlong, et al.
Published: (2026)
BAIT: Boundary-Guided Disclosure Escalation via Self-Conditioned Reasoning
by: Luo, Xuan, et al.
Published: (2026)
by: Luo, Xuan, et al.
Published: (2026)
Safeguarding Federated Learning-based Road Condition Classification
by: Liu, Sheng, et al.
Published: (2025)
by: Liu, Sheng, et al.
Published: (2025)
Proactive defense against LLM Jailbreak
by: Zhao, Weiliang, et al.
Published: (2025)
by: Zhao, Weiliang, et al.
Published: (2025)
FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks
by: Chen, Bocheng, et al.
Published: (2024)
by: Chen, Bocheng, et al.
Published: (2024)
Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection
by: Lin, Lixing, et al.
Published: (2026)
by: Lin, Lixing, et al.
Published: (2026)
Blind Spots in the Guard: How Domain-Camouflaged Injection Attacks Evade Detection in Multi-Agent LLM Systems
by: Pai, Aaditya
Published: (2026)
by: Pai, Aaditya
Published: (2026)
ThinkGuard: Deliberative Slow Thinking Leads to Cautious Guardrails
by: Wen, Xiaofei, et al.
Published: (2025)
by: Wen, Xiaofei, et al.
Published: (2025)
LLM Anonymization Against Agentic Re-Identification
by: Li, Ziwen, et al.
Published: (2026)
by: Li, Ziwen, et al.
Published: (2026)
Similar Items
-
GLiNER2-PII: A Multilingual Model for Personally Identifiable Information Extraction
by: Zaratiana, Urchade, et al.
Published: (2026) -
GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface
by: Zaratiana, Urchade, et al.
Published: (2025) -
Pioneer Agent: Continual Improvement of Small Language Models in Production
by: Atreja, Dhruv, et al.
Published: (2026) -
GLiNER Guard: Unified Encoder Family for Production LLM Safety and Privacy
by: Minko, Bogdan, et al.
Published: (2026) -
MaskSQL: Safeguarding Privacy for LLM-Based Text-to-SQL via Abstraction
by: Abedini, Sepideh, et al.
Published: (2025)