When Safety Becomes a Vulnerability: Exploiting LLM Alignment Homogeneity for Transferable Blocking in RAG
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Junchen, Qi, Chao, Wang, Rongzheng, Chen, Qizhi, Xu, Liang, Liang, Di, Simons, Bob, Liang, Shuang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KEPo: Knowledge Evolution Poison on Graph-based Retrieval-Augmented Generation
by: Chen, Qizhi, et al.
Published: (2026)
by: Chen, Qizhi, et al.
Published: (2026)
The Trust Paradox in LLM-Based Multi-Agent Systems: When Collaboration Becomes a Security Vulnerability
by: Xu, Zijie, et al.
Published: (2025)
by: Xu, Zijie, et al.
Published: (2025)
When Safe Models Merge into Danger: Exploiting Latent Vulnerabilities in LLM Fusion
by: Li, Jiaqing, et al.
Published: (2026)
by: Li, Jiaqing, et al.
Published: (2026)
When RSA Fails: Exploiting Prime Selection Vulnerabilities in Public Key Cryptography
by: Nikzad, Murtaza, et al.
Published: (2025)
by: Nikzad, Murtaza, et al.
Published: (2025)
Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
by: Li, Yuxi, et al.
Published: (2024)
by: Li, Yuxi, et al.
Published: (2024)
Colluding LoRA: A Compositional Vulnerability in LLM Safety Alignment
by: Ding, Sihao
Published: (2026)
by: Ding, Sihao
Published: (2026)
Exploiting Cross-Layer Vulnerabilities: Off-Path Attacks on the TCP/IP Protocol Suite
by: Feng, Xuewei, et al.
Published: (2024)
by: Feng, Xuewei, et al.
Published: (2024)
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms
by: He, Sinan, et al.
Published: (2025)
by: He, Sinan, et al.
Published: (2025)
When AIOps Become "AI Oops": Subverting LLM-driven IT Operations via Telemetry Manipulation
by: Pasquini, Dario, et al.
Published: (2025)
by: Pasquini, Dario, et al.
Published: (2025)
LLM Agents can Autonomously Exploit One-day Vulnerabilities
by: Fang, Richard, et al.
Published: (2024)
by: Fang, Richard, et al.
Published: (2024)
Don't Trust Your Upstream: Exploiting LLM Multi-Agent System via Topology-Guided Adversarial Propagation
by: Liang, Ruichao, et al.
Published: (2025)
by: Liang, Ruichao, et al.
Published: (2025)
RLCracker: Evaluating the Worst-Case Vulnerability of LLM Watermarks with Adaptive RL Attacks
by: Huang, Hanbo, et al.
Published: (2025)
by: Huang, Hanbo, et al.
Published: (2025)
The RAG Paradox: A Black-Box Attack Exploiting Unintentional Vulnerabilities in Retrieval-Augmented Generation Systems
by: Choi, Chanwoo, et al.
Published: (2025)
by: Choi, Chanwoo, et al.
Published: (2025)
Involuntary In-Context Learning: Exploiting Few-Shot Pattern Completion to Bypass Safety Alignment in GPT-5.4
by: Polyakov, Alex, et al.
Published: (2026)
by: Polyakov, Alex, et al.
Published: (2026)
A Game-Theoretic Foundation for Bitcoin's Price: A Security-Utility Equilibrium
by: Chen, Liang
Published: (2025)
by: Chen, Liang
Published: (2025)
KernJC: Automated Vulnerable Environment Generation for Linux Kernel Vulnerabilities
by: Ruan, Bonan, et al.
Published: (2024)
by: Ruan, Bonan, et al.
Published: (2024)
Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning
by: Wang, Yanbo, et al.
Published: (2026)
by: Wang, Yanbo, et al.
Published: (2026)
A Large-scale Empirical Study on the Generalizability of Disclosed Java Library Vulnerability Exploits
by: Chen, Zirui, et al.
Published: (2026)
by: Chen, Zirui, et al.
Published: (2026)
Breaking Precision Time: OS Vulnerability Exploits Against IEEE 1588
by: Soomro, Muhammad Abdullah, et al.
Published: (2025)
by: Soomro, Muhammad Abdullah, et al.
Published: (2025)
Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents
by: Nawal, Aditya, et al.
Published: (2026)
by: Nawal, Aditya, et al.
Published: (2026)
RASA: Routing-Aware Safety Alignment for Mixture-of-Experts Models
by: Liang, Jiacheng, et al.
Published: (2026)
by: Liang, Jiacheng, et al.
Published: (2026)
Safety Context Injection: Inference-Time Safety Alignment via Static Filtering and Agentic Analysis
by: Xu, Zhenhao, et al.
Published: (2026)
by: Xu, Zhenhao, et al.
Published: (2026)
HALURust: Exploiting Hallucinations of Large Language Models to Detect Vulnerabilities in Rust
by: Luo, Yu, et al.
Published: (2025)
by: Luo, Yu, et al.
Published: (2025)
CNT: Safety-oriented Function Reuse across LLMs via Cross-Model Neuron Transfer
by: Zhao, Yue, et al.
Published: (2026)
by: Zhao, Yue, et al.
Published: (2026)
Alleviating the Fear of Losing Alignment in LLM Fine-tuning
by: Yang, Kang, et al.
Published: (2025)
by: Yang, Kang, et al.
Published: (2025)
VIPER-MCP: Detecting and Exploiting Taint-Style Vulnerabilities in Model Context Protocol Servers
by: Sun, Pengyu, et al.
Published: (2026)
by: Sun, Pengyu, et al.
Published: (2026)
VulBinLLM: LLM-powered Vulnerability Detection for Stripped Binaries
by: Hussain, Nasir, et al.
Published: (2025)
by: Hussain, Nasir, et al.
Published: (2025)
Demystifying RCE Vulnerabilities in LLM-Integrated Apps
by: Liu, Tong, et al.
Published: (2023)
by: Liu, Tong, et al.
Published: (2023)
When Everyday Devices Become Weapons: A Closer Look at the Pager and Walkie-talkie Attacks
by: Sarker, Pantha Protim, et al.
Published: (2025)
by: Sarker, Pantha Protim, et al.
Published: (2025)
Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignment
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
Evil Vizier: Vulnerabilities of LLM-Integrated XR Systems
by: Zhang, Yicheng, et al.
Published: (2025)
by: Zhang, Yicheng, et al.
Published: (2025)
When Understanding Becomes a Risk: Authenticity and Safety Risks in the Emerging Image Generation Paradigm
by: Leng, Ye, et al.
Published: (2026)
by: Leng, Ye, et al.
Published: (2026)
SpatialJB: How Text Distribution Art Becomes the "Jailbreak Key" for LLM Guardrails
by: Mou, Zhiyi, et al.
Published: (2026)
by: Mou, Zhiyi, et al.
Published: (2026)
VulZoo: A Comprehensive Vulnerability Intelligence Dataset
by: Ruan, Bonan, et al.
Published: (2024)
by: Ruan, Bonan, et al.
Published: (2024)
When Convenience Becomes Risk: A Semantic View of Under-Specification in Host-Acting Agents
by: Lu, Di, et al.
Published: (2026)
by: Lu, Di, et al.
Published: (2026)
Select Me! When You Need a Tool: A Black-box Text Attack on Tool Selection
by: Chen, Liuji, et al.
Published: (2025)
by: Chen, Liuji, et al.
Published: (2025)
MCP Safety Training: Learning to Refuse Falsely Benign MCP Exploits using Improved Preference Alignment
by: Halloran, John
Published: (2025)
by: Halloran, John
Published: (2025)
Eradicating the Unseen: Detecting, Exploiting, and Remediating a Path Traversal Vulnerability across GitHub
by: Akhoundali, Jafar, et al.
Published: (2025)
by: Akhoundali, Jafar, et al.
Published: (2025)
When Developer Aid Becomes Security Debt: A Systematic Analysis of Insecure Behaviors in LLM Coding Agents
by: Kozak, Matous, et al.
Published: (2025)
by: Kozak, Matous, et al.
Published: (2025)
When Evaluation Becomes a Side Channel: Regime Leakage and Structural Mitigations for Alignment Assessment
by: Santos-Grueiro, Igor
Published: (2026)
by: Santos-Grueiro, Igor
Published: (2026)
Similar Items
-
KEPo: Knowledge Evolution Poison on Graph-based Retrieval-Augmented Generation
by: Chen, Qizhi, et al.
Published: (2026) -
The Trust Paradox in LLM-Based Multi-Agent Systems: When Collaboration Becomes a Security Vulnerability
by: Xu, Zijie, et al.
Published: (2025) -
When Safe Models Merge into Danger: Exploiting Latent Vulnerabilities in LLM Fusion
by: Li, Jiaqing, et al.
Published: (2026) -
When RSA Fails: Exploiting Prime Selection Vulnerabilities in Public Key Cryptography
by: Nikzad, Murtaza, et al.
Published: (2025) -
Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
by: Li, Yuxi, et al.
Published: (2024)