Probing Semantic Routing in Large Mixture-of-Expert Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913850536755200 |
|---|---|
| author | Olson, Matthew Lyle Ratzlaff, Neale Hinck, Musashi Luo, Man Yu, Sungduk Xue, Chendi Lal, Vasudev |
| author_facet | Olson, Matthew Lyle Ratzlaff, Neale Hinck, Musashi Luo, Man Yu, Sungduk Xue, Chendi Lal, Vasudev |
| contents | In the past year, large (>100B parameter) mixture-of-expert (MoE) models have become increasingly common in the open domain. While their advantages are often framed in terms of efficiency, prior work has also explored functional differentiation through routing behavior. We investigate whether expert routing in large MoE models is influenced by the semantics of the inputs. To test this, we design two controlled experiments. First, we compare activations on sentence pairs with a shared target word used in the same or different senses. Second, we fix context and substitute the target word with semantically similar or dissimilar alternatives. Comparing expert overlap across these conditions reveals clear, statistically significant evidence of semantic routing in large MoE models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_10928 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Probing Semantic Routing in Large Mixture-of-Expert Models Olson, Matthew Lyle Ratzlaff, Neale Hinck, Musashi Luo, Man Yu, Sungduk Xue, Chendi Lal, Vasudev Machine Learning Artificial Intelligence Computation and Language In the past year, large (>100B parameter) mixture-of-expert (MoE) models have become increasingly common in the open domain. While their advantages are often framed in terms of efficiency, prior work has also explored functional differentiation through routing behavior. We investigate whether expert routing in large MoE models is influenced by the semantics of the inputs. To test this, we design two controlled experiments. First, we compare activations on sentence pairs with a shared target word used in the same or different senses. Second, we fix context and substitute the target word with semantically similar or dissimilar alternatives. Comparing expert overlap across these conditions reveals clear, statistically significant evidence of semantic routing in large MoE models. |
| title | Probing Semantic Routing in Large Mixture-of-Expert Models |
| topic | Machine Learning Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2502.10928 |