Probing Semantic Routing in Large Mixture-of-Expert Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Olson, Matthew Lyle, Ratzlaff, Neale, Hinck, Musashi, Luo, Man, Yu, Sungduk, Xue, Chendi, Lal, Vasudev
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913850536755200
author Olson, Matthew Lyle
Ratzlaff, Neale
Hinck, Musashi
Luo, Man
Yu, Sungduk
Xue, Chendi
Lal, Vasudev
author_facet Olson, Matthew Lyle
Ratzlaff, Neale
Hinck, Musashi
Luo, Man
Yu, Sungduk
Xue, Chendi
Lal, Vasudev
contents In the past year, large (>100B parameter) mixture-of-expert (MoE) models have become increasingly common in the open domain. While their advantages are often framed in terms of efficiency, prior work has also explored functional differentiation through routing behavior. We investigate whether expert routing in large MoE models is influenced by the semantics of the inputs. To test this, we design two controlled experiments. First, we compare activations on sentence pairs with a shared target word used in the same or different senses. Second, we fix context and substitute the target word with semantically similar or dissimilar alternatives. Comparing expert overlap across these conditions reveals clear, statistically significant evidence of semantic routing in large MoE models.
format Preprint
id arxiv_https___arxiv_org_abs_2502_10928
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Probing Semantic Routing in Large Mixture-of-Expert Models
Olson, Matthew Lyle
Ratzlaff, Neale
Hinck, Musashi
Luo, Man
Yu, Sungduk
Xue, Chendi
Lal, Vasudev
Machine Learning
Artificial Intelligence
Computation and Language
In the past year, large (>100B parameter) mixture-of-expert (MoE) models have become increasingly common in the open domain. While their advantages are often framed in terms of efficiency, prior work has also explored functional differentiation through routing behavior. We investigate whether expert routing in large MoE models is influenced by the semantics of the inputs. To test this, we design two controlled experiments. First, we compare activations on sentence pairs with a shared target word used in the same or different senses. Second, we fix context and substitute the target word with semantically similar or dissimilar alternatives. Comparing expert overlap across these conditions reveals clear, statistically significant evidence of semantic routing in large MoE models.
title Probing Semantic Routing in Large Mixture-of-Expert Models
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2502.10928