The Art of Defending: A Systematic Evaluation and Analysis of LLM Defense Strategies on Safety and Over-Defensiveness
Fuente:
arXiv
Saved in:
| Main Authors: | Varshney, Neeraj, Dolin, Pavel, Seth, Agastya, Baral, Chitta |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Chaos with Keywords: Exposing Large Language Models Sycophantic Hallucination to Misleading Keywords and Evaluating Defense Strategies
by: RRV, Aswin, et al.
Published: (2024)
by: RRV, Aswin, et al.
Published: (2024)
LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models
by: Parmar, Mihir, et al.
Published: (2024)
by: Parmar, Mihir, et al.
Published: (2024)
Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models
by: Patel, Nisarg, et al.
Published: (2024)
by: Patel, Nisarg, et al.
Published: (2024)
Investigating and Addressing Hallucinations of LLMs in Tasks Involving Negation
by: Varshney, Neeraj, et al.
Published: (2024)
by: Varshney, Neeraj, et al.
Published: (2024)
Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey
by: Dong, Zhichen, et al.
Published: (2024)
by: Dong, Zhichen, et al.
Published: (2024)
Towards LogiGLUE: A Brief Survey and A Benchmark for Analyzing Logical Reasoning Capabilities of Language Models
by: Luo, Man, et al.
Published: (2023)
by: Luo, Man, et al.
Published: (2023)
UnSeenTimeQA: Time-Sensitive Question-Answering Beyond LLMs' Memorization
by: Uddin, Md Nayem, et al.
Published: (2024)
by: Uddin, Md Nayem, et al.
Published: (2024)
SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation
by: Liu, Xiaoze, et al.
Published: (2024)
by: Liu, Xiaoze, et al.
Published: (2024)
ActionReasoningBench: Reasoning about Actions with and without Ramification Constraints
by: Handa, Divij, et al.
Published: (2024)
by: Handa, Divij, et al.
Published: (2024)
Insights into Alignment: Evaluating DPO and its Variants Across Multiple Tasks
by: Saeidi, Amir, et al.
Published: (2024)
by: Saeidi, Amir, et al.
Published: (2024)
ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion Models
by: Patel, Maitreya, et al.
Published: (2023)
by: Patel, Maitreya, et al.
Published: (2023)
Pruning Strategies for Backdoor Defense in LLMs
by: Chapagain, Santosh, et al.
Published: (2025)
by: Chapagain, Santosh, et al.
Published: (2025)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
by: Zeng, Yifan, et al.
Published: (2024)
by: Zeng, Yifan, et al.
Published: (2024)
Explicit Reasoning Makes Better Judges: A Systematic Study on Accuracy, Efficiency, and Robustness
by: Jayarao, Pratik, et al.
Published: (2025)
by: Jayarao, Pratik, et al.
Published: (2025)
Cutting Through the Noise: Boosting LLM Performance on Math Word Problems
by: Anantheswaran, Ujjwala, et al.
Published: (2024)
by: Anantheswaran, Ujjwala, et al.
Published: (2024)
Model-Agnostic Lifelong LLM Safety via Externalized Attack-Defense Co-Evolution
by: Zhang, Xiaozhe, et al.
Published: (2026)
by: Zhang, Xiaozhe, et al.
Published: (2026)
Map&Make: Schema Guided Text to Table Generation
by: Ahuja, Naman, et al.
Published: (2025)
by: Ahuja, Naman, et al.
Published: (2025)
ViTaB-A: Evaluating Multimodal Large Language Models on Visual Table Attribution
by: Alqurnawi, Yahia, et al.
Published: (2026)
by: Alqurnawi, Yahia, et al.
Published: (2026)
Hypothesis Generation for Materials Discovery and Design Using Goal-Driven and Constraint-Guided LLM Agents
by: Kumbhar, Shrinidhi, et al.
Published: (2025)
by: Kumbhar, Shrinidhi, et al.
Published: (2025)
MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation
by: Jiang, Weisen, et al.
Published: (2025)
by: Jiang, Weisen, et al.
Published: (2025)
One Trigger Token Is Enough: A Defense Strategy for Balancing Safety and Usability in Large Language Models
by: Gu, Haoran, et al.
Published: (2025)
by: Gu, Haoran, et al.
Published: (2025)
Rethinking Information Synthesis in Multimodal Question Answering A Multi-Agent Perspective
by: Rajput, Krishna Singh, et al.
Published: (2025)
by: Rajput, Krishna Singh, et al.
Published: (2025)
DistillGuard: Evaluating Defenses Against LLM Knowledge Distillation
by: Jiang, Bo
Published: (2026)
by: Jiang, Bo
Published: (2026)
Defensive Dual Masking for Robust Adversarial Defense
by: Yang, Wangli, et al.
Published: (2024)
by: Yang, Wangli, et al.
Published: (2024)
Contextualized Privacy Defense for LLM Agents
by: Wen, Yule, et al.
Published: (2026)
by: Wen, Yule, et al.
Published: (2026)
Defenses & Enablers For Skill Injection Attacks on Terminal Based Agents
by: Fujinuma, Yoshinari, et al.
Published: (2026)
by: Fujinuma, Yoshinari, et al.
Published: (2026)
Cross-Task Defense: Instruction-Tuning LLMs for Content Safety
by: Fu, Yu, et al.
Published: (2024)
by: Fu, Yu, et al.
Published: (2024)
$λ$-ECLIPSE: Multi-Concept Personalized Text-to-Image Diffusion Models by Leveraging CLIP Latent Space
by: Patel, Maitreya, et al.
Published: (2024)
by: Patel, Maitreya, et al.
Published: (2024)
GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning
by: Siingh, Shikhhar, et al.
Published: (2025)
by: Siingh, Shikhhar, et al.
Published: (2025)
Triple Preference Optimization: Achieving Better Alignment using a Single Step Optimization
by: Saeidi, Amir, et al.
Published: (2024)
by: Saeidi, Amir, et al.
Published: (2024)
Code Mixologist : A Practitioner's Guide to Building Code-Mixed LLMs
by: Gupta, Himanshu, et al.
Published: (2026)
by: Gupta, Himanshu, et al.
Published: (2026)
EvoDefense: Co-Evolving Black-Box Defense with Large Language Models
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
Advancing Jailbreak Strategies: A Hybrid Approach to Exploiting LLM Vulnerabilities and Bypassing Modern Defenses
by: Ahmed, Mohamed, et al.
Published: (2025)
by: Ahmed, Mohamed, et al.
Published: (2025)
An Embarrassingly Simple Defense Against LLM Abliteration Attacks
by: Shairah, Harethah Abu, et al.
Published: (2025)
by: Shairah, Harethah Abu, et al.
Published: (2025)
From Recall to Forgetting: Benchmarking Long-Term Memory for Personalized Agents
by: Uddin, Md Nayem, et al.
Published: (2026)
by: Uddin, Md Nayem, et al.
Published: (2026)
Don't Blame the Annotator: Bias Already Starts in the Annotation Instructions
by: Parmar, Mihir, et al.
Published: (2022)
by: Parmar, Mihir, et al.
Published: (2022)
A Graph-Enhanced Defense Framework for Explainable Fake News Detection with LLM
by: Wang, Bo, et al.
Published: (2026)
by: Wang, Bo, et al.
Published: (2026)
AttentionDefense: Leveraging System Prompt Attention for Explainable Defense Against Novel Jailbreaks
by: Siska, Charlotte, et al.
Published: (2025)
by: Siska, Charlotte, et al.
Published: (2025)
A Systematic Literature Review on LLM Defenses Against Prompt Injection and Jailbreaking: Expanding NIST Taxonomy
by: Correia, Pedro H. Barcha, et al.
Published: (2026)
by: Correia, Pedro H. Barcha, et al.
Published: (2026)
LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
by: Li, Nathaniel, et al.
Published: (2024)
by: Li, Nathaniel, et al.
Published: (2024)
Similar Items
-
Chaos with Keywords: Exposing Large Language Models Sycophantic Hallucination to Misleading Keywords and Evaluating Defense Strategies
by: RRV, Aswin, et al.
Published: (2024) -
LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models
by: Parmar, Mihir, et al.
Published: (2024) -
Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models
by: Patel, Nisarg, et al.
Published: (2024) -
Investigating and Addressing Hallucinations of LLMs in Tasks Involving Negation
by: Varshney, Neeraj, et al.
Published: (2024) -
Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey
by: Dong, Zhichen, et al.
Published: (2024)