Learning How Much to Think: Difficulty-Aware Dynamic MoEs for Graph Node Classification

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhou, Jiajun, Li, Yadong, Chen, Xuanze, Ma, Chen, Zhao, Chuang, Yu, Shanqing, Xuan, Qi
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914469216518144
author Zhou, Jiajun
Li, Yadong
Chen, Xuanze
Ma, Chen
Zhao, Chuang
Yu, Shanqing
Xuan, Qi
author_facet Zhou, Jiajun
Li, Yadong
Chen, Xuanze
Ma, Chen
Zhao, Chuang
Yu, Shanqing
Xuan, Qi
contents Mixture-of-Experts (MoE) architectures offer a scalable path for Graph Neural Networks (GNNs) in node classification tasks but typically rely on static and rigid routing strategies that enforce a uniform expert budget or coarse-grained expert toggles on all nodes. This limitation overlooks the varying discriminative difficulty of nodes and leads to under-fitting for hard nodes and redundant computation for easy ones. To resolve this issue, we propose D2MoE, a novel framework that shifts the focus from static expert selection to node-wise expert resource allocation. By using predictive entropy as a real-time proxy for difficulty, D2MoE employs a difficulty-driven top-p routing mechanism to adaptively concentrate expert resources on hard nodes while reducing overhead for easy ones, achieving continuous and fine-grained expert budget scaling for node classification. Experiments on 13 benchmarks demonstrate that D2MoE achieves consistent state-of-the-art performance, surpassing leading baselines by up to 7.92% in accuracy on heterophilous graphs. Notably, on large-scale graphs, it reduces memory consumption by up to 73.07% and training time by 46.53% compared to the best-performing Graph MoE, thereby validating its superior efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2604_11473
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learning How Much to Think: Difficulty-Aware Dynamic MoEs for Graph Node Classification
Zhou, Jiajun
Li, Yadong
Chen, Xuanze
Ma, Chen
Zhao, Chuang
Yu, Shanqing
Xuan, Qi
Machine Learning
Mixture-of-Experts (MoE) architectures offer a scalable path for Graph Neural Networks (GNNs) in node classification tasks but typically rely on static and rigid routing strategies that enforce a uniform expert budget or coarse-grained expert toggles on all nodes. This limitation overlooks the varying discriminative difficulty of nodes and leads to under-fitting for hard nodes and redundant computation for easy ones. To resolve this issue, we propose D2MoE, a novel framework that shifts the focus from static expert selection to node-wise expert resource allocation. By using predictive entropy as a real-time proxy for difficulty, D2MoE employs a difficulty-driven top-p routing mechanism to adaptively concentrate expert resources on hard nodes while reducing overhead for easy ones, achieving continuous and fine-grained expert budget scaling for node classification. Experiments on 13 benchmarks demonstrate that D2MoE achieves consistent state-of-the-art performance, surpassing leading baselines by up to 7.92% in accuracy on heterophilous graphs. Notably, on large-scale graphs, it reduces memory consumption by up to 73.07% and training time by 46.53% compared to the best-performing Graph MoE, thereby validating its superior efficiency.
title Learning How Much to Think: Difficulty-Aware Dynamic MoEs for Graph Node Classification
topic Machine Learning
url https://arxiv.org/abs/2604.11473