Deconstructing Pre-training: Knowledge Attribution Analysis in MoE and Dense Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Bo, Li, Junzhuo, Chen, Hong, Chu, Yuanlin, Fan, Yuxuan, Hu, Xuming
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909988771856384
author Wang, Bo
Li, Junzhuo
Chen, Hong
Chu, Yuanlin
Fan, Yuxuan
Hu, Xuming
author_facet Wang, Bo
Li, Junzhuo
Chen, Hong
Chu, Yuanlin
Fan, Yuxuan
Hu, Xuming
contents Mixture-of-Experts (MoE) architectures decouple model capacity from per-token computation, enabling scaling beyond the computational limits imposed by dense scaling laws. Yet how MoE architectures shape knowledge acquisition during pre-training, and how this process differs from dense architectures, remains unknown. To address this issue, we introduce Gated-LPI (Log-Probability Increase), a neuron-level attribution metric that decomposes log-probability increase across neurons. We present a time-resolved comparison of knowledge acquisition dynamics in MoE and dense architectures, tracking checkpoints over 1.2M training steps (~ 5.0T tokens) and 600K training steps (~ 2.5T tokens), respectively. Our experiments uncover three patterns: (1) Low-entropy backbone. The top approximately 1% of MoE neurons capture over 45% of positive updates, forming a high-utility core, which is absent in the dense baseline. (2) Early consolidation. The MoE model locks into a stable importance profile within < 100K steps, whereas the dense model remains volatile throughout training. (3) Functional robustness. Masking the ten most important MoE attention heads reduces relational HIT@10 by < 10%, compared with > 50% for the dense model, showing that sparsity fosters distributed -- rather than brittle -- knowledge storage. These patterns collectively demonstrate that sparsity fosters an intrinsically stable and distributed computational backbone from early in training, helping bridge the gap between sparse architectures and training-time interpretability.
format Preprint
id arxiv_https___arxiv_org_abs_2601_08383
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Deconstructing Pre-training: Knowledge Attribution Analysis in MoE and Dense Models
Wang, Bo
Li, Junzhuo
Chen, Hong
Chu, Yuanlin
Fan, Yuxuan
Hu, Xuming
Artificial Intelligence
Computation and Language
Machine Learning
Mixture-of-Experts (MoE) architectures decouple model capacity from per-token computation, enabling scaling beyond the computational limits imposed by dense scaling laws. Yet how MoE architectures shape knowledge acquisition during pre-training, and how this process differs from dense architectures, remains unknown. To address this issue, we introduce Gated-LPI (Log-Probability Increase), a neuron-level attribution metric that decomposes log-probability increase across neurons. We present a time-resolved comparison of knowledge acquisition dynamics in MoE and dense architectures, tracking checkpoints over 1.2M training steps (~ 5.0T tokens) and 600K training steps (~ 2.5T tokens), respectively. Our experiments uncover three patterns: (1) Low-entropy backbone. The top approximately 1% of MoE neurons capture over 45% of positive updates, forming a high-utility core, which is absent in the dense baseline. (2) Early consolidation. The MoE model locks into a stable importance profile within < 100K steps, whereas the dense model remains volatile throughout training. (3) Functional robustness. Masking the ten most important MoE attention heads reduces relational HIT@10 by < 10%, compared with > 50% for the dense model, showing that sparsity fosters distributed -- rather than brittle -- knowledge storage. These patterns collectively demonstrate that sparsity fosters an intrinsically stable and distributed computational backbone from early in training, helping bridge the gap between sparse architectures and training-time interpretability.
title Deconstructing Pre-training: Knowledge Attribution Analysis in MoE and Dense Models
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2601.08383