MANATEE: Inference-Time Lightweight Diffusion Based Safety Defense for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Kan, Chun Yan Ryan, Tran, Tommy, Yadav, Vedant, Cai, Ava, Zhu, Kevin, Li, Ruizhe, Chaudhary, Maheep |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Neighborhood Blending: A Lightweight Inference-Time Defense Against Membership Inference Attacks
by: Zafar, Osama, et al.
Published: (2026)
by: Zafar, Osama, et al.
Published: (2026)
Zero-Shot Embedding Drift Detection: A Lightweight Defense Against Prompt Injections in LLMs
by: Sekar, Anirudh, et al.
Published: (2026)
by: Sekar, Anirudh, et al.
Published: (2026)
Weight space Detection of Backdoors in LoRA Adapters
by: Merenciano, David Puertolas, et al.
Published: (2026)
by: Merenciano, David Puertolas, et al.
Published: (2026)
Safety Context Injection: Inference-Time Safety Alignment via Static Filtering and Agentic Analysis
by: Xu, Zhenhao, et al.
Published: (2026)
by: Xu, Zhenhao, et al.
Published: (2026)
Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks
by: Kuo, Kevin, et al.
Published: (2026)
by: Kuo, Kevin, et al.
Published: (2026)
Cross-Task Defense: Instruction-Tuning LLMs for Content Safety
by: Fu, Yu, et al.
Published: (2024)
by: Fu, Yu, et al.
Published: (2024)
Evaluating Lightweight Block Cipher Payload Encryption for Real-Time CAN Traffic
by: Setterstrom, Kevin, et al.
Published: (2026)
by: Setterstrom, Kevin, et al.
Published: (2026)
Evolutionary Trigger Detection and Lightweight Model Repair Based Backdoor Defense
by: Zhou, Qi, et al.
Published: (2024)
by: Zhou, Qi, et al.
Published: (2024)
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
by: Xiong, Chen, et al.
Published: (2024)
by: Xiong, Chen, et al.
Published: (2024)
Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs
by: Yan, Dong, et al.
Published: (2026)
by: Yan, Dong, et al.
Published: (2026)
Botnet Detection on CTU-13 Using Lightweight Machine Learning Models
by: Gurappa, Subhash, et al.
Published: (2026)
by: Gurappa, Subhash, et al.
Published: (2026)
Plato's Form: Toward Backdoor Defense-as-a-Service for LLMs with Prototype Representations
by: Chen, Chen, et al.
Published: (2026)
by: Chen, Chen, et al.
Published: (2026)
Dynamic Probabilistic Noise Injection for Membership Inference Defense
by: Forough, Javad, et al.
Published: (2025)
by: Forough, Javad, et al.
Published: (2025)
Ensemble Privacy Defense for Knowledge-Intensive LLMs against Membership Inference Attacks
by: Fu, Haowei, et al.
Published: (2025)
by: Fu, Haowei, et al.
Published: (2025)
Jailbreaking LLMs & VLMs: Mechanisms, Evaluation, and Unified Defense
by: Chen, Zejian, et al.
Published: (2026)
by: Chen, Zejian, et al.
Published: (2026)
Efficient Adversarial Malware Defense via Trust-Based Raw Override and Confidence-Adaptive Bit-Depth Reduction
by: Chaudhary, Ayush, et al.
Published: (2025)
by: Chaudhary, Ayush, et al.
Published: (2025)
Membership Inference Attacks and Defenses in Federated Learning: A Survey
by: Bai, Li, et al.
Published: (2024)
by: Bai, Li, et al.
Published: (2024)
Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
by: Tong, Haibo, et al.
Published: (2025)
by: Tong, Haibo, et al.
Published: (2025)
Deep Learning Model Inversion Attacks and Defenses: A Comprehensive Survey
by: Yang, Wencheng, et al.
Published: (2025)
by: Yang, Wencheng, et al.
Published: (2025)
Proactive Hardening of LLM Defenses with HASTE
by: Chen, Henry, et al.
Published: (2026)
by: Chen, Henry, et al.
Published: (2026)
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense
by: Zhang, Jiawen, et al.
Published: (2025)
by: Zhang, Jiawen, et al.
Published: (2025)
Class-Conditional Neural Polarizer: A Lightweight and Effective Backdoor Defense by Purifying Poisoned Features
by: Zhu, Mingli, et al.
Published: (2025)
by: Zhu, Mingli, et al.
Published: (2025)
Minimal Cascade Gradient Smoothing for Fast Transferable Preemptive Adversarial Defense
by: Wang, Hanrui, et al.
Published: (2024)
by: Wang, Hanrui, et al.
Published: (2024)
Evaluating the Defense Potential of Machine Unlearning against Membership Inference Attacks
by: Tsiolakis, Theodoros, et al.
Published: (2025)
by: Tsiolakis, Theodoros, et al.
Published: (2025)
United We Defend: Collaborative Membership Inference Defenses in Federated Learning
by: Bai, Li, et al.
Published: (2026)
by: Bai, Li, et al.
Published: (2026)
From LLMs to MLLMs to Agents: A Survey of Emerging Paradigms in Jailbreak Attacks and Defenses within LLM Ecosystem
by: Mao, Yanxu, et al.
Published: (2025)
by: Mao, Yanxu, et al.
Published: (2025)
ICON: Indirect Prompt Injection Defense for Agents based on Inference-Time Correction
by: Wang, Che, et al.
Published: (2026)
by: Wang, Che, et al.
Published: (2026)
LightDefense: A Lightweight Uncertainty-Driven Defense against Jailbreaks via Shifted Token Distribution
by: Yang, Zhuoran, et al.
Published: (2025)
by: Yang, Zhuoran, et al.
Published: (2025)
Lite-BD: A Lightweight Black-box Backdoor Defense via Reviving Multi-Stage Image Transformations
by: Miah, Abdullah Arafat, et al.
Published: (2026)
by: Miah, Abdullah Arafat, et al.
Published: (2026)
Model-Agnostic Lifelong LLM Safety via Externalized Attack-Defense Co-Evolution
by: Zhang, Xiaozhe, et al.
Published: (2026)
by: Zhang, Xiaozhe, et al.
Published: (2026)
Adaptive Federated Learning Defences via Trust-Aware Deep Q-Networks
by: Palit, Vedant
Published: (2025)
by: Palit, Vedant
Published: (2025)
Black-box Membership Inference Attacks against Fine-tuned Diffusion Models
by: Pang, Yan, et al.
Published: (2023)
by: Pang, Yan, et al.
Published: (2023)
Detection and Defense Against Prominent Attacks on Preconditioned LLM-Integrated Virtual Assistants
by: Chan, Chun Fai, et al.
Published: (2024)
by: Chan, Chun Fai, et al.
Published: (2024)
Test-time Adversarial Defense with Opposite Adversarial Path and High Attack Time Cost
by: Yeh, Cheng-Han, et al.
Published: (2024)
by: Yeh, Cheng-Han, et al.
Published: (2024)
A Lightweight Defense Mechanism against Next Generation of Phishing Emails using Distilled Attention-Augmented BiLSTM
by: Eskandarian, Morteza, et al.
Published: (2026)
by: Eskandarian, Morteza, et al.
Published: (2026)
White-box Membership Inference Attacks against Diffusion Models
by: Pang, Yan, et al.
Published: (2023)
by: Pang, Yan, et al.
Published: (2023)
SALT: Steering Activations towards Leakage-free Thinking in Chain of Thought
by: Batra, Shourya, et al.
Published: (2025)
by: Batra, Shourya, et al.
Published: (2025)
On the (In-)Security of the Shuffling Defense in the Transformer Secure Inference
by: Li, Zhengyi, et al.
Published: (2026)
by: Li, Zhengyi, et al.
Published: (2026)
Time-Complexity Characterization of NIST Lightweight Cryptography Finalists
by: Hasan, Najmul, et al.
Published: (2026)
by: Hasan, Najmul, et al.
Published: (2026)
Targeting Alignment: Extracting Safety Classifiers of Aligned LLMs
by: Ferrand, Jean-Charles Noirot, et al.
Published: (2025)
by: Ferrand, Jean-Charles Noirot, et al.
Published: (2025)
Similar Items
-
Neighborhood Blending: A Lightweight Inference-Time Defense Against Membership Inference Attacks
by: Zafar, Osama, et al.
Published: (2026) -
Zero-Shot Embedding Drift Detection: A Lightweight Defense Against Prompt Injections in LLMs
by: Sekar, Anirudh, et al.
Published: (2026) -
Weight space Detection of Backdoors in LoRA Adapters
by: Merenciano, David Puertolas, et al.
Published: (2026) -
Safety Context Injection: Inference-Time Safety Alignment via Static Filtering and Agentic Analysis
by: Xu, Zhenhao, et al.
Published: (2026) -
Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks
by: Kuo, Kevin, et al.
Published: (2026)