Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Jinhwa, Harris, Ian G. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning
by: Wang, Yanbo, et al.
Published: (2026)
by: Wang, Yanbo, et al.
Published: (2026)
Medical Malice: A Dataset for Context-Aware Safety in Healthcare LLMs
by: D'addario, Andrew Maranhão Ventura
Published: (2025)
by: D'addario, Andrew Maranhão Ventura
Published: (2025)
PISanitizer: Preventing Prompt Injection to Long-Context LLMs via Prompt Sanitization
by: Geng, Runpeng, et al.
Published: (2025)
by: Geng, Runpeng, et al.
Published: (2025)
PrimeGuard: Safe and Helpful LLMs through Tuning-Free Routing
by: Manczak, Blazej, et al.
Published: (2024)
by: Manczak, Blazej, et al.
Published: (2024)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
by: Shao, Zedian, et al.
Published: (2024)
by: Shao, Zedian, et al.
Published: (2024)
Defend LLMs Through Self-Consciousness
by: Huang, Boshi, et al.
Published: (2025)
by: Huang, Boshi, et al.
Published: (2025)
Protecting Your LLMs with Information Bottleneck
by: Liu, Zichuan, et al.
Published: (2024)
by: Liu, Zichuan, et al.
Published: (2024)
Conversational Context Classification: A Representation Engineering Approach
by: Pan, Jonathan
Published: (2026)
by: Pan, Jonathan
Published: (2026)
SciSafeEval: A Comprehensive Benchmark for Safety Alignment of Large Language Models in Scientific Tasks
by: Li, Tianhao, et al.
Published: (2024)
by: Li, Tianhao, et al.
Published: (2024)
Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs
by: Xu, Zhao, et al.
Published: (2024)
by: Xu, Zhao, et al.
Published: (2024)
The Ethics of Interaction: Mitigating Security Threats in LLMs
by: Kumar, Ashutosh, et al.
Published: (2024)
by: Kumar, Ashutosh, et al.
Published: (2024)
Universal and Context-Independent Triggers for Precise Control of LLM Outputs
by: Liang, Jiashuo, et al.
Published: (2024)
by: Liang, Jiashuo, et al.
Published: (2024)
Strategic Deflection: Defending LLMs from Logit Manipulation
by: Rachidy, Yassine, et al.
Published: (2025)
by: Rachidy, Yassine, et al.
Published: (2025)
PARASITE: Conditional System Prompt Poisoning to Hijack LLMs
by: Pham, Viet, et al.
Published: (2025)
by: Pham, Viet, et al.
Published: (2025)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
by: Liu, Fan, et al.
Published: (2024)
by: Liu, Fan, et al.
Published: (2024)
CCJA: Context-Coherent Jailbreak Attack for Aligned Large Language Models
by: Zhou, Guanghao, et al.
Published: (2025)
by: Zhou, Guanghao, et al.
Published: (2025)
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
by: Mireshghallah, Niloofar, et al.
Published: (2023)
by: Mireshghallah, Niloofar, et al.
Published: (2023)
Graph of Attacks: Improved Black-Box and Interpretable Jailbreaks for LLMs
by: Akbar-Tajari, Mohammad, et al.
Published: (2025)
by: Akbar-Tajari, Mohammad, et al.
Published: (2025)
bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
by: Ji, Wence, et al.
Published: (2025)
by: Ji, Wence, et al.
Published: (2025)
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
by: Liu, Songyang, et al.
Published: (2025)
by: Liu, Songyang, et al.
Published: (2025)
Topology Matters: Measuring Memory Leakage in Multi-Agent LLMs
by: Liu, Jinbo, et al.
Published: (2025)
by: Liu, Jinbo, et al.
Published: (2025)
Feint and Attack: Attention-Based Strategies for Jailbreaking and Protecting LLMs
by: Pu, Rui, et al.
Published: (2024)
by: Pu, Rui, et al.
Published: (2024)
SeqAR: Jailbreak LLMs with Sequential Auto-Generated Characters
by: Yang, Yan, et al.
Published: (2024)
by: Yang, Yan, et al.
Published: (2024)
Waterfall: Framework for Robust and Scalable Text Watermarking and Provenance for LLMs
by: Lau, Gregory Kang Ruey, et al.
Published: (2024)
by: Lau, Gregory Kang Ruey, et al.
Published: (2024)
Sandwich attack: Multi-language Mixture Adaptive Attack on LLMs
by: Upadhayay, Bibek, et al.
Published: (2024)
by: Upadhayay, Bibek, et al.
Published: (2024)
Less is More: Sparse Watermarking in LLMs with Enhanced Text Quality
by: Hoang, Duy C., et al.
Published: (2024)
by: Hoang, Duy C., et al.
Published: (2024)
Rapid Optimization for Jailbreaking LLMs via Subconscious Exploitation and Echopraxia
by: Shen, Guangyu, et al.
Published: (2024)
by: Shen, Guangyu, et al.
Published: (2024)
Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
by: Hu, Xiaomeng, et al.
Published: (2025)
by: Hu, Xiaomeng, et al.
Published: (2025)
ShieldLearner: A New Paradigm for Jailbreak Attack Defense in LLMs
by: Ni, Ziyi, et al.
Published: (2025)
by: Ni, Ziyi, et al.
Published: (2025)
Effective and Efficient Jailbreaks of Black-Box LLMs with Cross-Behavior Attacks
by: Gohil, Vasudev
Published: (2025)
by: Gohil, Vasudev
Published: (2025)
Dagger Behind Smile: Fool LLMs with a Happy Ending Story
by: Song, Xurui, et al.
Published: (2025)
by: Song, Xurui, et al.
Published: (2025)
What Features in Prompts Jailbreak LLMs? Investigating the Mechanisms Behind Attacks
by: Kirch, Nathalie, et al.
Published: (2024)
by: Kirch, Nathalie, et al.
Published: (2024)
Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers
by: Wei, Jiali, et al.
Published: (2026)
by: Wei, Jiali, et al.
Published: (2026)
Swiss-Bench 003: Evaluating LLM Reliability and Adversarial Security for Swiss Regulatory Contexts
by: Uenal, Fatih
Published: (2026)
by: Uenal, Fatih
Published: (2026)
RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent
by: Xu, Huiyu, et al.
Published: (2024)
by: Xu, Huiyu, et al.
Published: (2024)
Injecting Falsehoods: Adversarial Man-in-the-Middle Attacks Undermining Factual Recall in LLMs
by: Fastowski, Alina, et al.
Published: (2025)
by: Fastowski, Alina, et al.
Published: (2025)
LoopLLM: Transferable Energy-Latency Attacks in LLMs via Repetitive Generation
by: Li, Xingyu, et al.
Published: (2025)
by: Li, Xingyu, et al.
Published: (2025)
DuFFin: A Dual-Level Fingerprinting Framework for LLMs IP Protection
by: Yan, Yuliang, et al.
Published: (2025)
by: Yan, Yuliang, et al.
Published: (2025)
From Vulnerabilities to Remediation: A Systematic Literature Review of LLMs in Code Security
by: Basic, Enna, et al.
Published: (2024)
by: Basic, Enna, et al.
Published: (2024)
Similar Items
-
Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning
by: Wang, Yanbo, et al.
Published: (2026) -
Medical Malice: A Dataset for Context-Aware Safety in Healthcare LLMs
by: D'addario, Andrew Maranhão Ventura
Published: (2025) -
PISanitizer: Preventing Prompt Injection to Long-Context LLMs via Prompt Sanitization
by: Geng, Runpeng, et al.
Published: (2025) -
PrimeGuard: Safe and Helpful LLMs through Tuning-Free Routing
by: Manczak, Blazej, et al.
Published: (2024) -
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
by: Shao, Zedian, et al.
Published: (2024)