Large Language Lobotomy: Jailbreaking Mixture-of-Experts via Expert Silencing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lintelo, Jona te, Wu, Lichao, Picek, Stjepan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BadPatches: Routing-aware Backdoor Attacks on Vision Mixture of Experts
von: Chan, Cedric, et al.
Veröffentlicht: (2025)
von: Chan, Cedric, et al.
Veröffentlicht: (2025)
MASCing: Configurable Mixture-of-Experts Behavior via Activation Steering Masks
von: Lintelo, Jona te, et al.
Veröffentlicht: (2026)
von: Lintelo, Jona te, et al.
Veröffentlicht: (2026)
The SkipSponge Attack: Sponge Weight Poisoning of Deep Neural Networks
von: Lintelo, Jona te, et al.
Veröffentlicht: (2024)
von: Lintelo, Jona te, et al.
Veröffentlicht: (2024)
GoodVibe: Security-by-Vibe for LLM-Based Code Generation
von: Thang, Maximilian, et al.
Veröffentlicht: (2026)
von: Thang, Maximilian, et al.
Veröffentlicht: (2026)
GateBreaker: Gate-Guided Attacks on Mixture-of-Expert LLMs
von: Wu, Lichao, et al.
Veröffentlicht: (2025)
von: Wu, Lichao, et al.
Veröffentlicht: (2025)
Backdoor Attacks on Decentralised Post-Training
von: Ersoy, Oğuzhan, et al.
Veröffentlicht: (2026)
von: Ersoy, Oğuzhan, et al.
Veröffentlicht: (2026)
A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models
von: Xu, Zihao, et al.
Veröffentlicht: (2024)
von: Xu, Zihao, et al.
Veröffentlicht: (2024)
Continuous Embedding Attacks via Clipped Inputs in Jailbreaking Large Language Models
von: Xu, Zihao, et al.
Veröffentlicht: (2024)
von: Xu, Zihao, et al.
Veröffentlicht: (2024)
$$\mathbf{L^2\cdot M = C^2}$$ Large Language Models are Covert Channels
von: Gaure, Simen, et al.
Veröffentlicht: (2024)
von: Gaure, Simen, et al.
Veröffentlicht: (2024)
NeuroStrike: Neuron-Level Attacks on Aligned LLMs
von: Wu, Lichao, et al.
Veröffentlicht: (2025)
von: Wu, Lichao, et al.
Veröffentlicht: (2025)
CatBack: Universal Backdoor Attacks on Tabular Data via Categorical Encoding
von: Tajalli, Behrad, et al.
Veröffentlicht: (2025)
von: Tajalli, Behrad, et al.
Veröffentlicht: (2025)
CryptoMoE: Privacy-Preserving and Scalable Mixture of Experts Inference via Balanced Expert Routing
von: Zhou, Yifan, et al.
Veröffentlicht: (2025)
von: Zhou, Yifan, et al.
Veröffentlicht: (2025)
Interpreting Emergent Features in Deep Learning-based Side-channel Analysis
von: Karayalçin, Sengim, et al.
Veröffentlicht: (2025)
von: Karayalçin, Sengim, et al.
Veröffentlicht: (2025)
EmoBack: Backdoor Attacks Against Speaker Identification Using Emotional Prosody
von: Schoof, Coen, et al.
Veröffentlicht: (2024)
von: Schoof, Coen, et al.
Veröffentlicht: (2024)
Who Speaks for the Trigger? Dynamic Expert Routing in Backdoored Mixture-of-Experts Transformers
von: Zhao, Xin, et al.
Veröffentlicht: (2025)
von: Zhao, Xin, et al.
Veröffentlicht: (2025)
Context is the Key: Backdoor Attacks for In-Context Learning with Vision Transformers
von: Abad, Gorka, et al.
Veröffentlicht: (2024)
von: Abad, Gorka, et al.
Veröffentlicht: (2024)
Flashy Backdoor: Real-world Environment Backdoor Attack on SNNs with DVS Cameras
von: Riaño, Roberto, et al.
Veröffentlicht: (2024)
von: Riaño, Roberto, et al.
Veröffentlicht: (2024)
Membership Privacy Evaluation in Deep Spiking Neural Networks
von: Li, Jiaxin, et al.
Veröffentlicht: (2024)
von: Li, Jiaxin, et al.
Veröffentlicht: (2024)
NoMod: A Non-modular Attack on Module Learning With Errors
von: Bassotto, Cristian, et al.
Veröffentlicht: (2025)
von: Bassotto, Cristian, et al.
Veröffentlicht: (2025)
Let's Focus: Focused Backdoor Attack against Federated Transfer Learning
von: Arazzi, Marco, et al.
Veröffentlicht: (2024)
von: Arazzi, Marco, et al.
Veröffentlicht: (2024)
Towards Backdoor Stealthiness in Model Parameter Space
von: Xu, Xiaoyun, et al.
Veröffentlicht: (2025)
von: Xu, Xiaoyun, et al.
Veröffentlicht: (2025)
BadMoE: Backdooring Mixture-of-Experts LLMs via Optimizing Routing Triggers and Infecting Dormant Experts
von: Wang, Qingyue, et al.
Veröffentlicht: (2025)
von: Wang, Qingyue, et al.
Veröffentlicht: (2025)
You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation
von: Arazzi, Marco, et al.
Veröffentlicht: (2026)
von: Arazzi, Marco, et al.
Veröffentlicht: (2026)
Buffer Overflow in Mixture of Experts
von: Hayes, Jamie, et al.
Veröffentlicht: (2024)
von: Hayes, Jamie, et al.
Veröffentlicht: (2024)
Toward Efficient Inference Attacks: Shadow Model Sharing via Mixture-of-Experts
von: Bai, Li, et al.
Veröffentlicht: (2025)
von: Bai, Li, et al.
Veröffentlicht: (2025)
MoJE: Mixture of Jailbreak Experts, Naive Tabular Classifiers as Guard for Prompt Attacks
von: Cornacchia, Giandomenico, et al.
Veröffentlicht: (2024)
von: Cornacchia, Giandomenico, et al.
Veröffentlicht: (2024)
A Systematic Study on the Design of Odd-Sized Highly Nonlinear Boolean Functions via Evolutionary Algorithms
von: Carlet, Claude, et al.
Veröffentlicht: (2025)
von: Carlet, Claude, et al.
Veröffentlicht: (2025)
Tighter Risk Bounds for Mixtures of Experts
von: Akretche, Wissam, et al.
Veröffentlicht: (2024)
von: Akretche, Wissam, et al.
Veröffentlicht: (2024)
BAN: Detecting Backdoors Activated by Adversarial Neuron Noise
von: Xu, Xiaoyun, et al.
Veröffentlicht: (2024)
von: Xu, Xiaoyun, et al.
Veröffentlicht: (2024)
MoPE: A Mixture of Password Experts for Improving Password Guessing
von: Duan, Mingjian, et al.
Veröffentlicht: (2025)
von: Duan, Mingjian, et al.
Veröffentlicht: (2025)
Removing the Trigger, Not the Backdoor: Alternative Triggers and Latent Backdoors
von: Abad, Gorka, et al.
Veröffentlicht: (2026)
von: Abad, Gorka, et al.
Veröffentlicht: (2026)
Time-Distributed Backdoor Attacks on Federated Spiking Learning
von: Abad, Gorka, et al.
Veröffentlicht: (2024)
von: Abad, Gorka, et al.
Veröffentlicht: (2024)
Differentially Private Training of Mixture of Experts Models
von: Tholoniat, Pierre, et al.
Veröffentlicht: (2024)
von: Tholoniat, Pierre, et al.
Veröffentlicht: (2024)
Improving LLM Outputs Against Jailbreak Attacks with Expert Model Integration
von: Tsmindashvili, Tatia, et al.
Veröffentlicht: (2025)
von: Tsmindashvili, Tatia, et al.
Veröffentlicht: (2025)
The Power of Bamboo: On the Post-Compromise Security for Searchable Symmetric Encryption
von: Chen, Tianyang, et al.
Veröffentlicht: (2024)
von: Chen, Tianyang, et al.
Veröffentlicht: (2024)
Misrouter: Exploiting Routing Mechanisms for Input-Only Attacks on Mixture-of-Experts LLMs
von: Fei, Zekun, et al.
Veröffentlicht: (2026)
von: Fei, Zekun, et al.
Veröffentlicht: (2026)
Enhancing Physical Layer Communication Security through Generative AI with Mixture of Experts
von: Zhao, Changyuan, et al.
Veröffentlicht: (2024)
von: Zhao, Changyuan, et al.
Veröffentlicht: (2024)
Virtual Context: Enhancing Jailbreak Attacks with Special Token Injection
von: Zhou, Yuqi, et al.
Veröffentlicht: (2024)
von: Zhou, Yuqi, et al.
Veröffentlicht: (2024)
IDEM Enough? Evolving Highly Nonlinear Idempotent Boolean Functions
von: Carlet, Claude, et al.
Veröffentlicht: (2026)
von: Carlet, Claude, et al.
Veröffentlicht: (2026)
Degree is Important: On Evolving Homogeneous Boolean Functions
von: Carlet, Claude, et al.
Veröffentlicht: (2025)
von: Carlet, Claude, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
BadPatches: Routing-aware Backdoor Attacks on Vision Mixture of Experts
von: Chan, Cedric, et al.
Veröffentlicht: (2025) -
MASCing: Configurable Mixture-of-Experts Behavior via Activation Steering Masks
von: Lintelo, Jona te, et al.
Veröffentlicht: (2026) -
The SkipSponge Attack: Sponge Weight Poisoning of Deep Neural Networks
von: Lintelo, Jona te, et al.
Veröffentlicht: (2024) -
GoodVibe: Security-by-Vibe for LLM-Based Code Generation
von: Thang, Maximilian, et al.
Veröffentlicht: (2026) -
GateBreaker: Gate-Guided Attacks on Mixture-of-Expert LLMs
von: Wu, Lichao, et al.
Veröffentlicht: (2025)