Saved in:
Bibliographic Details
Main Authors: Kim, Jinhwa, Harris, Ian G.
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2508.10031
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916898007941120
author Kim, Jinhwa
Harris, Ian G.
author_facet Kim, Jinhwa
Harris, Ian G.
contents While Large Language Models (LLMs) have shown significant advancements in performance, various jailbreak attacks have posed growing safety and ethical risks. Malicious users often exploit adversarial context to deceive LLMs, prompting them to generate responses to harmful queries. In this study, we propose a new defense mechanism called Context Filtering model, an input pre-processing method designed to filter out untrustworthy and unreliable context while identifying the primary prompts containing the real user intent to uncover concealed malicious intent. Given that enhancing the safety of LLMs often compromises their helpfulness, potentially affecting the experience of benign users, our method aims to improve the safety of the LLMs while preserving their original performance. We evaluate the effectiveness of our model in defending against jailbreak attacks through comparative analysis, comparing our approach with state-of-the-art defense mechanisms against six different attacks and assessing the helpfulness of LLMs under these defenses. Our model demonstrates its ability to reduce the Attack Success Rates of jailbreak attacks by up to 88% while maintaining the original LLMs' performance, achieving state-of-the-art Safety and Helpfulness Product results. Notably, our model is a plug-and-play method that can be applied to all LLMs, including both white-box and black-box models, to enhance their safety without requiring any fine-tuning of the models themselves. We will make our model publicly available for research purposes.
format Preprint
id arxiv_https___arxiv_org_abs_2508_10031
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs
Kim, Jinhwa
Harris, Ian G.
Cryptography and Security
Artificial Intelligence
Computation and Language
While Large Language Models (LLMs) have shown significant advancements in performance, various jailbreak attacks have posed growing safety and ethical risks. Malicious users often exploit adversarial context to deceive LLMs, prompting them to generate responses to harmful queries. In this study, we propose a new defense mechanism called Context Filtering model, an input pre-processing method designed to filter out untrustworthy and unreliable context while identifying the primary prompts containing the real user intent to uncover concealed malicious intent. Given that enhancing the safety of LLMs often compromises their helpfulness, potentially affecting the experience of benign users, our method aims to improve the safety of the LLMs while preserving their original performance. We evaluate the effectiveness of our model in defending against jailbreak attacks through comparative analysis, comparing our approach with state-of-the-art defense mechanisms against six different attacks and assessing the helpfulness of LLMs under these defenses. Our model demonstrates its ability to reduce the Attack Success Rates of jailbreak attacks by up to 88% while maintaining the original LLMs' performance, achieving state-of-the-art Safety and Helpfulness Product results. Notably, our model is a plug-and-play method that can be applied to all LLMs, including both white-box and black-box models, to enhance their safety without requiring any fine-tuning of the models themselves. We will make our model publicly available for research purposes.
title Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs
topic Cryptography and Security
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2508.10031