Exploring and Developing a Pre-Model Safeguard with Draft Models
Fuente:
arXiv
Saved in:
| Main Authors: | Cai, Hongyu, Arunasalam, Arjun, Liang, Yiming, Bianchi, Antonio, Celik, Z. Berkay |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring and Developing a Pre-Model Safeguard with Draft Models
by: Cai, Hongyu, et al.
Published: (2026)
by: Cai, Hongyu, et al.
Published: (2026)
Rethinking How to Evaluate Language Model Jailbreak
by: Cai, Hongyu, et al.
Published: (2024)
by: Cai, Hongyu, et al.
Published: (2024)
International Students and Scams: At Risk Abroad
by: Zhang, Katherine, et al.
Published: (2025)
by: Zhang, Katherine, et al.
Published: (2025)
Enhancing LLM-based Autonomous Driving Agents to Mitigate Perception Attacks
by: Song, Ruoyu, et al.
Published: (2024)
by: Song, Ruoyu, et al.
Published: (2024)
Investigating the Impact of Dark Patterns on LLM-Based Web Agents
by: Ersoy, Devin, et al.
Published: (2025)
by: Ersoy, Devin, et al.
Published: (2025)
Formalizing the Safety, Security, and Functional Properties of Agentic AI Systems
by: Allegrini, Edoardo, et al.
Published: (2025)
by: Allegrini, Edoardo, et al.
Published: (2025)
Safeguarding Large Language Models: A Survey
by: Dong, Yi, et al.
Published: (2024)
by: Dong, Yi, et al.
Published: (2024)
Safeguarding AI Agents: Developing and Analyzing Safety Architectures
by: Domkundwar, Ishaan, et al.
Published: (2024)
by: Domkundwar, Ishaan, et al.
Published: (2024)
RoboJailBench: Benchmarking Adversarial Attacks and Defenses in Embodied Robotic Agents
by: Yeke, Doguhuan, et al.
Published: (2026)
by: Yeke, Doguhuan, et al.
Published: (2026)
LM-Scout: Analyzing the Security of Language Model Integration in Android Apps
by: Ibrahim, Muhammad, et al.
Published: (2025)
by: Ibrahim, Muhammad, et al.
Published: (2025)
TransLinkGuard: Safeguarding Transformer Models Against Model Stealing in Edge Deployment
by: Li, Qinfeng, et al.
Published: (2024)
by: Li, Qinfeng, et al.
Published: (2024)
LLM Safeguard is a Double-Edged Sword: Exploiting False Positives for Denial-of-Service Attacks
by: Zhang, Qingzhao, et al.
Published: (2024)
by: Zhang, Qingzhao, et al.
Published: (2024)
Safeguarding Text-to-Image Generative Models Against Unauthorized Knowledge Distillation
by: Gao, Yilan, et al.
Published: (2026)
by: Gao, Yilan, et al.
Published: (2026)
Embedding with Large Language Models for Classification of HIPAA Safeguard Compliance Rules
by: Rahman, Md Abdur, et al.
Published: (2024)
by: Rahman, Md Abdur, et al.
Published: (2024)
Do Not Merge My Model! Safeguarding Open-Source LLMs Against Unauthorized Model Merging
by: Li, Qinfeng, et al.
Published: (2025)
by: Li, Qinfeng, et al.
Published: (2025)
Delayed Backdoor Attacks: Exploring the Temporal Dimension as a New Attack Surface in Pre-Trained Models
by: Ding, Zikang, et al.
Published: (2026)
by: Ding, Zikang, et al.
Published: (2026)
Shadows in the Code: Exploring the Risks and Defenses of LLM-based Multi-Agent Software Development Systems
by: Wang, Xiaoqing, et al.
Published: (2025)
by: Wang, Xiaoqing, et al.
Published: (2025)
On Evaluating the Durability of Safeguards for Open-Weight LLMs
by: Qi, Xiangyu, et al.
Published: (2024)
by: Qi, Xiangyu, et al.
Published: (2024)
PromptKeeper: Safeguarding System Prompts for LLMs
by: Jiang, Zhifeng, et al.
Published: (2024)
by: Jiang, Zhifeng, et al.
Published: (2024)
Safeguarding Multimodal Knowledge Copyright in the RAG-as-a-Service Environment
by: Chen, Tianyu, et al.
Published: (2025)
by: Chen, Tianyu, et al.
Published: (2025)
AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting
by: Wang, Yu, et al.
Published: (2024)
by: Wang, Yu, et al.
Published: (2024)
Re-Triggering Safeguards within LLMs for Jailbreak Detection
by: Lin, Zheng, et al.
Published: (2026)
by: Lin, Zheng, et al.
Published: (2026)
Safeguarding Federated Learning-based Road Condition Classification
by: Liu, Sheng, et al.
Published: (2025)
by: Liu, Sheng, et al.
Published: (2025)
Reliable Model Watermarking: Defending Against Theft without Compromising on Evasion
by: Zhu, Hongyu, et al.
Published: (2024)
by: Zhu, Hongyu, et al.
Published: (2024)
GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
by: Liu, Yue, et al.
Published: (2025)
by: Liu, Yue, et al.
Published: (2025)
TERD: A Unified Framework for Safeguarding Diffusion Models Against Backdoors
by: Mo, Yichuan, et al.
Published: (2024)
by: Mo, Yichuan, et al.
Published: (2024)
Persistent Pre-Training Poisoning of LLMs
by: Zhang, Yiming, et al.
Published: (2024)
by: Zhang, Yiming, et al.
Published: (2024)
CrossGuard: Safeguarding MLLMs against Joint-Modal Implicit Malicious Attacks
by: Zhang, Xu, et al.
Published: (2025)
by: Zhang, Xu, et al.
Published: (2025)
GuardReasoner: Towards Reasoning-based LLM Safeguards
by: Liu, Yue, et al.
Published: (2025)
by: Liu, Yue, et al.
Published: (2025)
Shattering the Echo Chamber: Hidden Safeguards in Manuscripts Against the AI Takeover of Peer Review
by: Ma, Oubo, et al.
Published: (2026)
by: Ma, Oubo, et al.
Published: (2026)
MCP Guardian: A Security-First Layer for Safeguarding MCP-Based AI System
by: Kumar, Sonu, et al.
Published: (2025)
by: Kumar, Sonu, et al.
Published: (2025)
CoopGuard: Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Round Attacks
by: Li, Siyuan, et al.
Published: (2026)
by: Li, Siyuan, et al.
Published: (2026)
Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection
by: Lin, Lixing, et al.
Published: (2026)
by: Lin, Lixing, et al.
Published: (2026)
SentinelNet: Safeguarding Multi-Agent Collaboration Through Credit-Based Dynamic Threat Detection
by: Feng, Yang, et al.
Published: (2025)
by: Feng, Yang, et al.
Published: (2025)
Understanding Users' Security and Privacy Concerns and Attitudes Towards Conversational AI Platforms
by: Ali, Mutahar, et al.
Published: (2025)
by: Ali, Mutahar, et al.
Published: (2025)
CORVUS: Red-Teaming Hallucination Detectors via Internal Signal Camouflage in Large Language Models
by: Min, Nay Myat, et al.
Published: (2026)
by: Min, Nay Myat, et al.
Published: (2026)
Membership Inference for Contrastive Pre-training Models with Text-only PII Queries
by: Cheng, Ruoxi, et al.
Published: (2026)
by: Cheng, Ruoxi, et al.
Published: (2026)
DSSmoothing: Toward Certified Dataset Ownership Verification for Pre-trained Language Models via Dual-Space Smoothing
by: Qiao, Ting, et al.
Published: (2025)
by: Qiao, Ting, et al.
Published: (2025)
MAGE: Safeguarding LLM Agents against Long-Horizon Threats via Shadow Memory
by: Wang, Yuhui, et al.
Published: (2026)
by: Wang, Yuhui, et al.
Published: (2026)
Backdoor Samples Detection Based on Perturbation Discrepancy Consistency in Pre-trained Language Models
by: Peng, Zuquan, et al.
Published: (2025)
by: Peng, Zuquan, et al.
Published: (2025)
Similar Items
-
Exploring and Developing a Pre-Model Safeguard with Draft Models
by: Cai, Hongyu, et al.
Published: (2026) -
Rethinking How to Evaluate Language Model Jailbreak
by: Cai, Hongyu, et al.
Published: (2024) -
International Students and Scams: At Risk Abroad
by: Zhang, Katherine, et al.
Published: (2025) -
Enhancing LLM-based Autonomous Driving Agents to Mitigate Perception Attacks
by: Song, Ruoyu, et al.
Published: (2024) -
Investigating the Impact of Dark Patterns on LLM-Based Web Agents
by: Ersoy, Devin, et al.
Published: (2025)