MASCing: Configurable Mixture-of-Experts Behavior via Activation Steering Masks
Fuente:
arXiv
Saved in:
| Main Authors: | Lintelo, Jona te, Wu, Lichao, Krček, Marina, Karayalçin, Sengim, Picek, Stjepan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Large Language Lobotomy: Jailbreaking Mixture-of-Experts via Expert Silencing
by: Lintelo, Jona te, et al.
Published: (2026)
by: Lintelo, Jona te, et al.
Published: (2026)
Interpreting Emergent Features in Deep Learning-based Side-channel Analysis
by: Karayalçin, Sengim, et al.
Published: (2025)
by: Karayalçin, Sengim, et al.
Published: (2025)
BadPatches: Routing-aware Backdoor Attacks on Vision Mixture of Experts
by: Chan, Cedric, et al.
Published: (2025)
by: Chan, Cedric, et al.
Published: (2025)
Backdoor Directions in Vision Transformers
by: Karayalcin, Sengim, et al.
Published: (2026)
by: Karayalcin, Sengim, et al.
Published: (2026)
The SkipSponge Attack: Sponge Weight Poisoning of Deep Neural Networks
by: Lintelo, Jona te, et al.
Published: (2024)
by: Lintelo, Jona te, et al.
Published: (2024)
Backdoor Attacks on Decentralised Post-Training
by: Ersoy, Oğuzhan, et al.
Published: (2026)
by: Ersoy, Oğuzhan, et al.
Published: (2026)
GoodVibe: Security-by-Vibe for LLM-Based Code Generation
by: Thang, Maximilian, et al.
Published: (2026)
by: Thang, Maximilian, et al.
Published: (2026)
GateBreaker: Gate-Guided Attacks on Mixture-of-Expert LLMs
by: Wu, Lichao, et al.
Published: (2025)
by: Wu, Lichao, et al.
Published: (2025)
NoMod: A Non-modular Attack on Module Learning With Errors
by: Bassotto, Cristian, et al.
Published: (2025)
by: Bassotto, Cristian, et al.
Published: (2025)
Label Inference Attacks against Node-level Vertical Federated GNNs
by: Arazzi, Marco, et al.
Published: (2023)
by: Arazzi, Marco, et al.
Published: (2023)
NeuroStrike: Neuron-Level Attacks on Aligned LLMs
by: Wu, Lichao, et al.
Published: (2025)
by: Wu, Lichao, et al.
Published: (2025)
CatBack: Universal Backdoor Attacks on Tabular Data via Categorical Encoding
by: Tajalli, Behrad, et al.
Published: (2025)
by: Tajalli, Behrad, et al.
Published: (2025)
SoK: The Last Line of Defense: On Backdoor Defense Evaluation
by: Abad, Gorka, et al.
Published: (2025)
by: Abad, Gorka, et al.
Published: (2025)
BAN: Detecting Backdoors Activated by Adversarial Neuron Noise
by: Xu, Xiaoyun, et al.
Published: (2024)
by: Xu, Xiaoyun, et al.
Published: (2024)
$$\mathbf{L^2\cdot M = C^2}$$ Large Language Models are Covert Channels
by: Gaure, Simen, et al.
Published: (2024)
by: Gaure, Simen, et al.
Published: (2024)
EmoBack: Backdoor Attacks Against Speaker Identification Using Emotional Prosody
by: Schoof, Coen, et al.
Published: (2024)
by: Schoof, Coen, et al.
Published: (2024)
Context is the Key: Backdoor Attacks for In-Context Learning with Vision Transformers
by: Abad, Gorka, et al.
Published: (2024)
by: Abad, Gorka, et al.
Published: (2024)
Flashy Backdoor: Real-world Environment Backdoor Attack on SNNs with DVS Cameras
by: Riaño, Roberto, et al.
Published: (2024)
by: Riaño, Roberto, et al.
Published: (2024)
Membership Privacy Evaluation in Deep Spiking Neural Networks
by: Li, Jiaxin, et al.
Published: (2024)
by: Li, Jiaxin, et al.
Published: (2024)
Let's Focus: Focused Backdoor Attack against Federated Transfer Learning
by: Arazzi, Marco, et al.
Published: (2024)
by: Arazzi, Marco, et al.
Published: (2024)
Towards Backdoor Stealthiness in Model Parameter Space
by: Xu, Xiaoyun, et al.
Published: (2025)
by: Xu, Xiaoyun, et al.
Published: (2025)
You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation
by: Arazzi, Marco, et al.
Published: (2026)
by: Arazzi, Marco, et al.
Published: (2026)
A Systematic Study on the Design of Odd-Sized Highly Nonlinear Boolean Functions via Evolutionary Algorithms
by: Carlet, Claude, et al.
Published: (2025)
by: Carlet, Claude, et al.
Published: (2025)
A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models
by: Xu, Zihao, et al.
Published: (2024)
by: Xu, Zihao, et al.
Published: (2024)
Removing the Trigger, Not the Backdoor: Alternative Triggers and Latent Backdoors
by: Abad, Gorka, et al.
Published: (2026)
by: Abad, Gorka, et al.
Published: (2026)
Time-Distributed Backdoor Attacks on Federated Spiking Learning
by: Abad, Gorka, et al.
Published: (2024)
by: Abad, Gorka, et al.
Published: (2024)
The Power of Bamboo: On the Post-Compromise Security for Searchable Symmetric Encryption
by: Chen, Tianyang, et al.
Published: (2024)
by: Chen, Tianyang, et al.
Published: (2024)
IDEM Enough? Evolving Highly Nonlinear Idempotent Boolean Functions
by: Carlet, Claude, et al.
Published: (2026)
by: Carlet, Claude, et al.
Published: (2026)
Degree is Important: On Evolving Homogeneous Boolean Functions
by: Carlet, Claude, et al.
Published: (2025)
by: Carlet, Claude, et al.
Published: (2025)
A Systematic Evaluation of Evolving Highly Nonlinear Boolean Functions in Odd Sizes
by: Carlet, Claude, et al.
Published: (2024)
by: Carlet, Claude, et al.
Published: (2024)
CryptoMoE: Privacy-Preserving and Scalable Mixture of Experts Inference via Balanced Expert Routing
by: Zhou, Yifan, et al.
Published: (2025)
by: Zhou, Yifan, et al.
Published: (2025)
Sneaky Spikes: Uncovering Stealthy Backdoor Attacks in Spiking Neural Networks with Neuromorphic Data
by: Abad, Gorka, et al.
Published: (2023)
by: Abad, Gorka, et al.
Published: (2023)
More is Better (Mostly): On the Backdoor Attacks in Federated Graph Neural Networks
by: Xu, Jing, et al.
Published: (2022)
by: Xu, Jing, et al.
Published: (2022)
Monotone but Exciting: On Evolving Monotone Boolean Functions with High Nonlinearity
by: Carlet, Claude, et al.
Published: (2026)
by: Carlet, Claude, et al.
Published: (2026)
NegaBent, No Regrets: Evolving Spectrally Flat Boolean Functions
by: Carlet, Claude, et al.
Published: (2026)
by: Carlet, Claude, et al.
Published: (2026)
On Counts and Densities of Homogeneous Bent Functions: An Evolutionary Approach
by: Carlet, Claude, et al.
Published: (2025)
by: Carlet, Claude, et al.
Published: (2025)
Backdoor Attacks on Transformers for Tabular Data: An Empirical Study
by: Pleiter, Bart, et al.
Published: (2023)
by: Pleiter, Bart, et al.
Published: (2023)
PII Jailbreaking in LLMs via Activation Steering Reveals Personal Information Leakage
by: Nakka, Krishna Kanth, et al.
Published: (2025)
by: Nakka, Krishna Kanth, et al.
Published: (2025)
Continuous Embedding Attacks via Clipped Inputs in Jailbreaking Large Language Models
by: Xu, Zihao, et al.
Published: (2024)
by: Xu, Zihao, et al.
Published: (2024)
Started Off Local, Now We're in the Cloud: Forensic Examination of the Amazon Echo Show 15 Smart Display
by: Crasselt, Jona, et al.
Published: (2024)
by: Crasselt, Jona, et al.
Published: (2024)
Similar Items
-
Large Language Lobotomy: Jailbreaking Mixture-of-Experts via Expert Silencing
by: Lintelo, Jona te, et al.
Published: (2026) -
Interpreting Emergent Features in Deep Learning-based Side-channel Analysis
by: Karayalçin, Sengim, et al.
Published: (2025) -
BadPatches: Routing-aware Backdoor Attacks on Vision Mixture of Experts
by: Chan, Cedric, et al.
Published: (2025) -
Backdoor Directions in Vision Transformers
by: Karayalcin, Sengim, et al.
Published: (2026) -
The SkipSponge Attack: Sponge Weight Poisoning of Deep Neural Networks
by: Lintelo, Jona te, et al.
Published: (2024)