BadMoE: Backdooring Mixture-of-Experts LLMs via Optimizing Routing Triggers and Infecting Dormant Experts
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Qingyue, Pang, Qi, Lin, Xixun, Wang, Shuai, Wu, Daoyuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Who Speaks for the Trigger? Dynamic Expert Routing in Backdoored Mixture-of-Experts Transformers
by: Zhao, Xin, et al.
Published: (2025)
by: Zhao, Xin, et al.
Published: (2025)
BadPatches: Routing-aware Backdoor Attacks on Vision Mixture of Experts
by: Chan, Cedric, et al.
Published: (2025)
by: Chan, Cedric, et al.
Published: (2025)
CryptoMoE: Privacy-Preserving and Scalable Mixture of Experts Inference via Balanced Expert Routing
by: Zhou, Yifan, et al.
Published: (2025)
by: Zhou, Yifan, et al.
Published: (2025)
Misrouter: Exploiting Routing Mechanisms for Input-Only Attacks on Mixture-of-Experts LLMs
by: Fei, Zekun, et al.
Published: (2026)
by: Fei, Zekun, et al.
Published: (2026)
Condition-Triggered Cryptographic Asset Control via Dormant Authorization Paths
by: Wang, Jian Sheng
Published: (2026)
by: Wang, Jian Sheng
Published: (2026)
Large Language Lobotomy: Jailbreaking Mixture-of-Experts via Expert Silencing
by: Lintelo, Jona te, et al.
Published: (2026)
by: Lintelo, Jona te, et al.
Published: (2026)
Backdoor Contrastive Learning via Bi-level Trigger Optimization
by: Sun, Weiyu, et al.
Published: (2024)
by: Sun, Weiyu, et al.
Published: (2024)
GateBreaker: Gate-Guided Attacks on Mixture-of-Expert LLMs
by: Wu, Lichao, et al.
Published: (2025)
by: Wu, Lichao, et al.
Published: (2025)
MoPE: A Mixture of Password Experts for Improving Password Guessing
by: Duan, Mingjian, et al.
Published: (2025)
by: Duan, Mingjian, et al.
Published: (2025)
RASA: Routing-Aware Safety Alignment for Mixture-of-Experts Models
by: Liang, Jiacheng, et al.
Published: (2026)
by: Liang, Jiacheng, et al.
Published: (2026)
Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
by: Jiang, Yukun, et al.
Published: (2026)
by: Jiang, Yukun, et al.
Published: (2026)
MoEcho: Exploiting Side-Channel Attacks to Compromise User Privacy in Mixture-of-Experts LLMs
by: Ding, Ruyi, et al.
Published: (2025)
by: Ding, Ruyi, et al.
Published: (2025)
TooBadRL: Trigger Optimization to Boost Effectiveness of Backdoor Attacks on Deep Reinforcement Learning
by: Zhang, Mingxuan, et al.
Published: (2025)
by: Zhang, Mingxuan, et al.
Published: (2025)
Differentiation-Based Extraction of Proprietary Data from Fine-Tuned LLMs
by: Li, Zongjie, et al.
Published: (2025)
by: Li, Zongjie, et al.
Published: (2025)
MASCing: Configurable Mixture-of-Experts Behavior via Activation Steering Masks
by: Lintelo, Jona te, et al.
Published: (2026)
by: Lintelo, Jona te, et al.
Published: (2026)
LLMs Can Defend Themselves Against Jailbreaking in a Practical Manner: A Vision Paper
by: Wu, Daoyuan, et al.
Published: (2024)
by: Wu, Daoyuan, et al.
Published: (2024)
Routing-Aware Explanations for Mixture of Experts Graph Models in Malware Detection
by: Shokouhinejad, Hossein, et al.
Published: (2026)
by: Shokouhinejad, Hossein, et al.
Published: (2026)
BadLLM-TG: A Backdoor Defender powered by LLM Trigger Generator
by: Zhang, Ruyi, et al.
Published: (2026)
by: Zhang, Ruyi, et al.
Published: (2026)
Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs
by: Wang, Yifei, et al.
Published: (2026)
by: Wang, Yifei, et al.
Published: (2026)
Backdoors in Code Summarizers: How Bad Is It?
by: Wang, Chenyu, et al.
Published: (2025)
by: Wang, Chenyu, et al.
Published: (2025)
Mixture of Robust Experts (MoRE):A Robust Denoising Method towards multiple perturbations
by: Cheng, Hao, et al.
Published: (2021)
by: Cheng, Hao, et al.
Published: (2021)
Isolate Trigger: Detecting and Eliminating Adaptive Backdoor Attacks
by: Sun, Chengrui, et al.
Published: (2025)
by: Sun, Chengrui, et al.
Published: (2025)
Buffer Overflow in Mixture of Experts
by: Hayes, Jamie, et al.
Published: (2024)
by: Hayes, Jamie, et al.
Published: (2024)
Toward Efficient Inference Attacks: Shadow Model Sharing via Mixture-of-Experts
by: Bai, Li, et al.
Published: (2025)
by: Bai, Li, et al.
Published: (2025)
SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert Identification
by: Lai, Zhenglin, et al.
Published: (2025)
by: Lai, Zhenglin, et al.
Published: (2025)
Meme Trojan: Backdoor Attacks Against Hateful Meme Detection via Cross-Modal Triggers
by: Wang, Ruofei, et al.
Published: (2024)
by: Wang, Ruofei, et al.
Published: (2024)
BadTemplate: A Training-Free Backdoor Attack via Chat Template Against Large Language Models
by: Wang, Zihan, et al.
Published: (2026)
by: Wang, Zihan, et al.
Published: (2026)
Tighter Risk Bounds for Mixtures of Experts
by: Akretche, Wissam, et al.
Published: (2024)
by: Akretche, Wissam, et al.
Published: (2024)
Revisiting Training-Inference Trigger Intensity in Backdoor Attacks
by: Lin, Chenhao, et al.
Published: (2025)
by: Lin, Chenhao, et al.
Published: (2025)
BadFair: Backdoored Fairness Attacks with Group-conditioned Triggers
by: Xue, Jiaqi, et al.
Published: (2024)
by: Xue, Jiaqi, et al.
Published: (2024)
RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry
by: Lv, Bo, et al.
Published: (2026)
by: Lv, Bo, et al.
Published: (2026)
BadDLM: Backdooring Diffusion Language Models with Diverse Targets
by: Zhai, Shengfang, et al.
Published: (2026)
by: Zhai, Shengfang, et al.
Published: (2026)
A Practical Trigger-Free Backdoor Attack on Neural Networks
by: Wang, Jiahao, et al.
Published: (2024)
by: Wang, Jiahao, et al.
Published: (2024)
BadRSSD: Backdoor Attacks on Regularized Self-Supervised Diffusion Models
by: Wang, Jiayao, et al.
Published: (2026)
by: Wang, Jiayao, et al.
Published: (2026)
BadTime: An Effective Backdoor Attack on Multivariate Long-Term Time Series Forecasting
by: Xiang, Kunlan, et al.
Published: (2025)
by: Xiang, Kunlan, et al.
Published: (2025)
TrafficMoE: Heterogeneity-aware Mixture of Experts for Encrypted Traffic Classification
by: He, Qing, et al.
Published: (2026)
by: He, Qing, et al.
Published: (2026)
Differentially Private Training of Mixture of Experts Models
by: Tholoniat, Pierre, et al.
Published: (2024)
by: Tholoniat, Pierre, et al.
Published: (2024)
Trigger Where It Hurts: Unveiling Hidden Backdoors through Sensitivity with Sensitron
by: Zhao, Gejian, et al.
Published: (2025)
by: Zhao, Gejian, et al.
Published: (2025)
Enhancing Physical Layer Communication Security through Generative AI with Mixture of Experts
by: Zhao, Changyuan, et al.
Published: (2024)
by: Zhao, Changyuan, et al.
Published: (2024)
BadEdit: Backdooring large language models by model editing
by: Li, Yanzhou, et al.
Published: (2024)
by: Li, Yanzhou, et al.
Published: (2024)
Similar Items
-
Who Speaks for the Trigger? Dynamic Expert Routing in Backdoored Mixture-of-Experts Transformers
by: Zhao, Xin, et al.
Published: (2025) -
BadPatches: Routing-aware Backdoor Attacks on Vision Mixture of Experts
by: Chan, Cedric, et al.
Published: (2025) -
CryptoMoE: Privacy-Preserving and Scalable Mixture of Experts Inference via Balanced Expert Routing
by: Zhou, Yifan, et al.
Published: (2025) -
Misrouter: Exploiting Routing Mechanisms for Input-Only Attacks on Mixture-of-Experts LLMs
by: Fei, Zekun, et al.
Published: (2026) -
Condition-Triggered Cryptographic Asset Control via Dormant Authorization Paths
by: Wang, Jian Sheng
Published: (2026)