When Safe Models Merge into Danger: Exploiting Latent Vulnerabilities in LLM Fusion
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Jiaqing, Zhang, Zhibo, Zhou, Shide, Li, Yuxi, Yu, Tianlong, Wang, Kailong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exposing the Ghost in the Transformer: Abnormal Detection for Large Language Models via Hidden State Forensics
by: Zhou, Shide, et al.
Published: (2025)
by: Zhou, Shide, et al.
Published: (2025)
Model-Editing-Based Jailbreak against Safety-aligned Large Language Models
by: Li, Yuxi, et al.
Published: (2024)
by: Li, Yuxi, et al.
Published: (2024)
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift
by: Yuan, Shuai, et al.
Published: (2025)
by: Yuan, Shuai, et al.
Published: (2025)
Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
by: Li, Yuxi, et al.
Published: (2024)
by: Li, Yuxi, et al.
Published: (2024)
When Safety Becomes a Vulnerability: Exploiting LLM Alignment Homogeneity for Transferable Blocking in RAG
by: Li, Junchen, et al.
Published: (2026)
by: Li, Junchen, et al.
Published: (2026)
"Glue pizza and eat rocks" -- Exploiting Vulnerabilities in Retrieval-Augmented Generative Models
by: Tan, Zhen, et al.
Published: (2024)
by: Tan, Zhen, et al.
Published: (2024)
When RSA Fails: Exploiting Prime Selection Vulnerabilities in Public Key Cryptography
by: Nikzad, Murtaza, et al.
Published: (2025)
by: Nikzad, Murtaza, et al.
Published: (2025)
Yama: Precise Opcode-based Data Flow Analysis for Detecting PHP Applications Vulnerabilities
by: Jiazhen, Zhao, et al.
Published: (2024)
by: Jiazhen, Zhao, et al.
Published: (2024)
HALURust: Exploiting Hallucinations of Large Language Models to Detect Vulnerabilities in Rust
by: Luo, Yu, et al.
Published: (2025)
by: Luo, Yu, et al.
Published: (2025)
UNSEEN: A Cross-Stack LLM Unlearning Defense against AR-LLM Social Engineering Attacks
by: Yu, Tianlong, et al.
Published: (2026)
by: Yu, Tianlong, et al.
Published: (2026)
VIPER-MCP: Detecting and Exploiting Taint-Style Vulnerabilities in Model Context Protocol Servers
by: Sun, Pengyu, et al.
Published: (2026)
by: Sun, Pengyu, et al.
Published: (2026)
The Trust Paradox in LLM-Based Multi-Agent Systems: When Collaboration Becomes a Security Vulnerability
by: Xu, Zijie, et al.
Published: (2025)
by: Xu, Zijie, et al.
Published: (2025)
LLM Agents can Autonomously Exploit One-day Vulnerabilities
by: Fang, Richard, et al.
Published: (2024)
by: Fang, Richard, et al.
Published: (2024)
Exploiting Cross-Layer Vulnerabilities: Off-Path Attacks on the TCP/IP Protocol Suite
by: Feng, Xuewei, et al.
Published: (2024)
by: Feng, Xuewei, et al.
Published: (2024)
Compiled Models, Built-In Exploits: Uncovering Pervasive Bit-Flip Attack Surfaces in DNN Executables
by: Chen, Yanzuo, et al.
Published: (2023)
by: Chen, Yanzuo, et al.
Published: (2023)
Enhancing Smart Contract Vulnerability Detection in DApps Leveraging Fine-Tuned LLM
by: Bu, Jiuyang, et al.
Published: (2025)
by: Bu, Jiuyang, et al.
Published: (2025)
The Vulnerability of LLM Rankers to Prompt Injection Attacks
by: Yin, Yu, et al.
Published: (2026)
by: Yin, Yu, et al.
Published: (2026)
Merge Hijacking: Backdoor Attacks to Model Merging of Large Language Models
by: Yuan, Zenghui, et al.
Published: (2025)
by: Yuan, Zenghui, et al.
Published: (2025)
STAC: When Innocent Tools Form Dangerous Chains to Jailbreak LLM Agents
by: Li, Jing-Jing, et al.
Published: (2025)
by: Li, Jing-Jing, et al.
Published: (2025)
RefineRAG: Word-Level Poisoning Attacks via Retriever-Guided Text Refinement
by: Wang, Ziye, et al.
Published: (2026)
by: Wang, Ziye, et al.
Published: (2026)
LLM-SmartAudit: Advanced Smart Contract Vulnerability Detection
by: Wei, Zhiyuan, et al.
Published: (2024)
by: Wei, Zhiyuan, et al.
Published: (2024)
Breaking Precision Time: OS Vulnerability Exploits Against IEEE 1588
by: Soomro, Muhammad Abdullah, et al.
Published: (2025)
by: Soomro, Muhammad Abdullah, et al.
Published: (2025)
SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code
by: Li, Xinghang, et al.
Published: (2025)
by: Li, Xinghang, et al.
Published: (2025)
LLM-BSCVM: An LLM-Based Blockchain Smart Contract Vulnerability Management Framework
by: Jin, Yanli, et al.
Published: (2025)
by: Jin, Yanli, et al.
Published: (2025)
CKG-LLM: LLM-Assisted Detection of Smart Contract Access Control Vulnerabilities Based on Knowledge Graphs
by: Li, Xiaoqi, et al.
Published: (2025)
by: Li, Xiaoqi, et al.
Published: (2025)
Vulnerability-Aware Robust Multimodal Adversarial Training
by: Zhang, Junrui, et al.
Published: (2025)
by: Zhang, Junrui, et al.
Published: (2025)
MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots
by: Deng, Gelei, et al.
Published: (2023)
by: Deng, Gelei, et al.
Published: (2023)
Demystifying RCE Vulnerabilities in LLM-Integrated Apps
by: Liu, Tong, et al.
Published: (2023)
by: Liu, Tong, et al.
Published: (2023)
WALLETRADAR: Towards Automating the Detection of Vulnerabilities in Browser-based Cryptocurrency Wallets
by: Xia, Pengcheng, et al.
Published: (2024)
by: Xia, Pengcheng, et al.
Published: (2024)
Neuro-symbolic Static Analysis with LLM-generated Vulnerability Patterns
by: Li, Penghui, et al.
Published: (2025)
by: Li, Penghui, et al.
Published: (2025)
The Shadow of Fraud: The Emerging Danger of AI-powered Social Engineering and its Possible Cure
by: Yu, Jingru, et al.
Published: (2024)
by: Yu, Jingru, et al.
Published: (2024)
BadMerging: Backdoor Attacks Against Model Merging
by: Zhang, Jinghuai, et al.
Published: (2024)
by: Zhang, Jinghuai, et al.
Published: (2024)
A Model for Assessing Network Asset Vulnerability Using QPSO-LightGBM
by: Li, Xinyu, et al.
Published: (2024)
by: Li, Xinyu, et al.
Published: (2024)
VULSOLVER: Vulnerability Detection via LLM-Driven Constraint Solving
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
RerouteGuard: Understanding and Mitigating Adversarial Risks for LLM Routing
by: Zhang, Wenhui, et al.
Published: (2026)
by: Zhang, Wenhui, et al.
Published: (2026)
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms
by: He, Sinan, et al.
Published: (2025)
by: He, Sinan, et al.
Published: (2025)
Evaluating and Enhancing the Vulnerability Reasoning Capabilities of Large Language Models
by: Lu, Li, et al.
Published: (2026)
by: Lu, Li, et al.
Published: (2026)
VulBinLLM: LLM-powered Vulnerability Detection for Stripped Binaries
by: Hussain, Nasir, et al.
Published: (2025)
by: Hussain, Nasir, et al.
Published: (2025)
Pandora: Jailbreak GPTs by Retrieval Augmented Generation Poisoning
by: Deng, Gelei, et al.
Published: (2024)
by: Deng, Gelei, et al.
Published: (2024)
The Danger Within: Insider Threat Modeling Using Business Process Models
by: von der Assen, Jan, et al.
Published: (2024)
by: von der Assen, Jan, et al.
Published: (2024)
Similar Items
-
Exposing the Ghost in the Transformer: Abnormal Detection for Large Language Models via Hidden State Forensics
by: Zhou, Shide, et al.
Published: (2025) -
Model-Editing-Based Jailbreak against Safety-aligned Large Language Models
by: Li, Yuxi, et al.
Published: (2024) -
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift
by: Yuan, Shuai, et al.
Published: (2025) -
Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
by: Li, Yuxi, et al.
Published: (2024) -
When Safety Becomes a Vulnerability: Exploiting LLM Alignment Homogeneity for Transferable Blocking in RAG
by: Li, Junchen, et al.
Published: (2026)