Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Yukun, Huang, Hai, Li, Mingjie, Zhang, Yage, Backes, Michael, Zhang, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RASA: Routing-Aware Safety Alignment for Mixture-of-Experts Models
by: Liang, Jiacheng, et al.
Published: (2026)
by: Liang, Jiacheng, et al.
Published: (2026)
HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?
by: Jiang, Yukun, et al.
Published: (2026)
by: Jiang, Yukun, et al.
Published: (2026)
Real Money, Fake Models: Deceptive Model Claims in Shadow APIs
by: Zhang, Yage, et al.
Published: (2026)
by: Zhang, Yage, et al.
Published: (2026)
Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities
by: Qu, Yiting, et al.
Published: (2025)
by: Qu, Yiting, et al.
Published: (2025)
"Humans welcome to observe": A First Look at the Agent Social Network Moltbook
by: Jiang, Yukun, et al.
Published: (2026)
by: Jiang, Yukun, et al.
Published: (2026)
Unsafe LLM-Based Search: Quantitative Analysis and Mitigation of Safety Risks in AI Web Search
by: Luo, Zeren, et al.
Published: (2025)
by: Luo, Zeren, et al.
Published: (2025)
Excessive Reasoning Attack on Reasoning LLMs
by: Si, Wai Man, et al.
Published: (2025)
by: Si, Wai Man, et al.
Published: (2025)
PM-MOE: Mixture of Experts on Private Model Parameters for Personalized Federated Learning
by: Feng, Yu, et al.
Published: (2025)
by: Feng, Yu, et al.
Published: (2025)
SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert Identification
by: Lai, Zhenglin, et al.
Published: (2025)
by: Lai, Zhenglin, et al.
Published: (2025)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
by: Chu, Junjie, et al.
Published: (2024)
by: Chu, Junjie, et al.
Published: (2024)
Reconstruct Your Previous Conversations! Comprehensively Investigating Privacy Leakage Risks in Conversations with GPT Models
by: Chu, Junjie, et al.
Published: (2024)
by: Chu, Junjie, et al.
Published: (2024)
Undetectable Backdoors in Model Parameters: Hiding Sparse Secrets in High Dimensions
by: Choudhary, Sarthak, et al.
Published: (2026)
by: Choudhary, Sarthak, et al.
Published: (2026)
TERD: A Unified Framework for Safeguarding Diffusion Models Against Backdoors
by: Mo, Yichuan, et al.
Published: (2024)
by: Mo, Yichuan, et al.
Published: (2024)
Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning
by: Liu, Guozhi, et al.
Published: (2025)
by: Liu, Guozhi, et al.
Published: (2025)
MoJE: Mixture of Jailbreak Experts, Naive Tabular Classifiers as Guard for Prompt Attacks
by: Cornacchia, Giandomenico, et al.
Published: (2024)
by: Cornacchia, Giandomenico, et al.
Published: (2024)
TeleSparse: Practical Privacy-Preserving Verification of Deep Neural Networks
by: Maheri, Mohammad M, et al.
Published: (2025)
by: Maheri, Mohammad M, et al.
Published: (2025)
Training on Fake Labels: Mitigating Label Leakage in Split Learning via Secure Dimension Transformation
by: Jiang, Yukun, et al.
Published: (2024)
by: Jiang, Yukun, et al.
Published: (2024)
Mixture of Robust Experts (MoRE):A Robust Denoising Method towards multiple perturbations
by: Cheng, Hao, et al.
Published: (2021)
by: Cheng, Hao, et al.
Published: (2021)
Efficient Privacy-Preserving Recommendation on Sparse Data using Fully Homomorphic Encryption
by: Chowdhury, Moontaha Nishat, et al.
Published: (2025)
by: Chowdhury, Moontaha Nishat, et al.
Published: (2025)
Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task Concurrency
by: Jiang, Yukun, et al.
Published: (2025)
by: Jiang, Yukun, et al.
Published: (2025)
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets
by: Lu, Ning, et al.
Published: (2025)
by: Lu, Ning, et al.
Published: (2025)
Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs
by: Wang, Yifei, et al.
Published: (2026)
by: Wang, Yifei, et al.
Published: (2026)
MCP Safety Audit: LLMs with the Model Context Protocol Allow Major Security Exploits
by: Radosevich, Brandon, et al.
Published: (2025)
by: Radosevich, Brandon, et al.
Published: (2025)
Guiding not Forcing: Enhancing the Transferability of Jailbreaking Attacks on LLMs via Removing Superfluous Constraints
by: Yang, Junxiao, et al.
Published: (2025)
by: Yang, Junxiao, et al.
Published: (2025)
UpSafe$^\circ$C: Upcycling for Controllable Safety in Large Language Models
by: Sun, Yuhao, et al.
Published: (2025)
by: Sun, Yuhao, et al.
Published: (2025)
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
by: Muhamed, Aashiq, et al.
Published: (2025)
by: Muhamed, Aashiq, et al.
Published: (2025)
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
BadMoE: Backdooring Mixture-of-Experts LLMs via Optimizing Routing Triggers and Infecting Dormant Experts
by: Wang, Qingyue, et al.
Published: (2025)
by: Wang, Qingyue, et al.
Published: (2025)
Stealing User Prompts from Mixture of Experts
by: Yona, Itay, et al.
Published: (2024)
by: Yona, Itay, et al.
Published: (2024)
Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable
by: Huang, Tiansheng, et al.
Published: (2025)
by: Huang, Tiansheng, et al.
Published: (2025)
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
MetaMoE: Diversity-Aware Proxy Selection for Privacy-Preserving Mixture-of-Experts Unification
by: Jiang, Weisen, et al.
Published: (2026)
by: Jiang, Weisen, et al.
Published: (2026)
Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
by: Panfilov, Alexander, et al.
Published: (2025)
by: Panfilov, Alexander, et al.
Published: (2025)
Image-Perfect Imperfections: Safety, Bias, and Authenticity in the Shadow of Text-To-Image Model Evolution
by: Wu, Yixin, et al.
Published: (2024)
by: Wu, Yixin, et al.
Published: (2024)
Towards Understanding the Robustness of Sparse Autoencoders
by: Saiyed, Ahson, et al.
Published: (2026)
by: Saiyed, Ahson, et al.
Published: (2026)
Routing-Aware Explanations for Mixture of Experts Graph Models in Malware Detection
by: Shokouhinejad, Hossein, et al.
Published: (2026)
by: Shokouhinejad, Hossein, et al.
Published: (2026)
SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection
by: Chen, Shuhao, et al.
Published: (2026)
by: Chen, Shuhao, et al.
Published: (2026)
Fast Exact Unlearning for In-Context Learning Data for LLMs
by: Muresanu, Andrei I., et al.
Published: (2024)
by: Muresanu, Andrei I., et al.
Published: (2024)
Uncovering, Explaining, and Mitigating the Superficial Safety of Backdoor Defense
by: Min, Rui, et al.
Published: (2024)
by: Min, Rui, et al.
Published: (2024)
LoBAM: LoRA-Based Backdoor Attack on Model Merging
by: Yin, Ming, et al.
Published: (2024)
by: Yin, Ming, et al.
Published: (2024)
Similar Items
-
RASA: Routing-Aware Safety Alignment for Mixture-of-Experts Models
by: Liang, Jiacheng, et al.
Published: (2026) -
HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?
by: Jiang, Yukun, et al.
Published: (2026) -
Real Money, Fake Models: Deceptive Model Claims in Shadow APIs
by: Zhang, Yage, et al.
Published: (2026) -
Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities
by: Qu, Yiting, et al.
Published: (2025) -
"Humans welcome to observe": A First Look at the Agent Social Network Moltbook
by: Jiang, Yukun, et al.
Published: (2026)