Expert Selections In MoE Models Reveal (Almost) As Much As Text

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Nuriyev, Amir, Kulp, Gabriel
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918385869127680
author Nuriyev, Amir
Kulp, Gabriel
author_facet Nuriyev, Amir
Kulp, Gabriel
contents We present a text-reconstruction attack on mixture-of-experts (MoE) language models that recovers tokens from expert selections alone. In MoE models, each token is routed to a subset of expert subnetworks; we show these routing decisions leak substantially more information than previously understood. Prior work using logistic regression achieves limited reconstruction; we show that a 3-layer MLP improves this to 63.1% top-1 accuracy, and that a transformer-based sequence decoder recovers 91.2% of tokens top-1 (94.8% top-10) on 32-token sequences from OpenWebText after training on 100M tokens. These results connect MoE routing to the broader literature on embedding inversion. We outline practical leakage scenarios (e.g., distributed inference and side channels) and show that adding noise reduces but does not eliminate reconstruction. Our findings suggest that expert selections in MoE deployments should be treated as sensitive as the underlying text.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04105
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Expert Selections In MoE Models Reveal (Almost) As Much As Text
Nuriyev, Amir
Kulp, Gabriel
Computation and Language
Cryptography and Security
We present a text-reconstruction attack on mixture-of-experts (MoE) language models that recovers tokens from expert selections alone. In MoE models, each token is routed to a subset of expert subnetworks; we show these routing decisions leak substantially more information than previously understood. Prior work using logistic regression achieves limited reconstruction; we show that a 3-layer MLP improves this to 63.1% top-1 accuracy, and that a transformer-based sequence decoder recovers 91.2% of tokens top-1 (94.8% top-10) on 32-token sequences from OpenWebText after training on 100M tokens. These results connect MoE routing to the broader literature on embedding inversion. We outline practical leakage scenarios (e.g., distributed inference and side channels) and show that adding noise reduces but does not eliminate reconstruction. Our findings suggest that expert selections in MoE deployments should be treated as sensitive as the underlying text.
title Expert Selections In MoE Models Reveal (Almost) As Much As Text
topic Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2602.04105