One Router to Route Them All: Homogeneous Expert Routing for Heterogeneous Graph Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Shakirov, Georgiy, Arakelov, Albert |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Task-Conditioned Routing Signatures in Sparse Mixture-of-Experts Transformers
by: Avinash, Mynampati Sri Ranganadha
Published: (2026)
by: Avinash, Mynampati Sri Ranganadha
Published: (2026)
ReflexGrad: Within-Episode Failure Recovery in LLM Agents via Progress-Gated Dual-Process Routing
by: Kadu, Ankush, et al.
Published: (2025)
by: Kadu, Ankush, et al.
Published: (2025)
Deep Memory Search: A Metaheuristic Approach for Optimizing Heuristic Search
by: Hedar, Abdel-Rahman, et al.
Published: (2024)
by: Hedar, Abdel-Rahman, et al.
Published: (2024)
Hidden Failure Modes of Gradient Modification under Adam in Continual Learning, and Adaptive Decoupled Moment Routing as a Repair
by: Hu, Yuelin, et al.
Published: (2026)
by: Hu, Yuelin, et al.
Published: (2026)
Not All Transitions Matter: Evidence from PPO
by: Basnet, Ajhesh
Published: (2026)
by: Basnet, Ajhesh
Published: (2026)
PolyGLU: State-Conditional Activation Routing in Transformer Feed-Forward Networks
by: Medeiros, Daniel Nobrega
Published: (2026)
by: Medeiros, Daniel Nobrega
Published: (2026)
A Geometric Perspective for High-Dimensional Multiplex Graphs
by: Abdous, Kamel, et al.
Published: (2025)
by: Abdous, Kamel, et al.
Published: (2025)
Graph Memory Transformer (GMT)
by: Zanarini, Nicola, et al.
Published: (2026)
by: Zanarini, Nicola, et al.
Published: (2026)
Superposition Is Not Necessary: A Mechanistic Interpretability Analysis of Transformer Representations for Time Series Forecasting
by: Yıldırım, Alper
Published: (2026)
by: Yıldırım, Alper
Published: (2026)
DRFormer: Multi-Scale Transformer Utilizing Diverse Receptive Fields for Long Time-Series Forecasting
by: Ding, Ruixin, et al.
Published: (2024)
by: Ding, Ruixin, et al.
Published: (2024)
Low-Dimensional Execution Manifolds in Transformer Learning Dynamics: Evidence from Modular Arithmetic Tasks
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
Simulation-Driven Railway Delay Prediction: An Imitation Learning Approach
by: Elliker, Clément, et al.
Published: (2025)
by: Elliker, Clément, et al.
Published: (2025)
Leveraging Personalized PageRank and Higher-Order Topological Structures for Heterophily Mitigation in Graph Neural Networks
by: Wang, Yumeng, et al.
Published: (2025)
by: Wang, Yumeng, et al.
Published: (2025)
Introducing New Node Prediction in Graph Mining: Predicting All Links from Isolated Nodes with Graph Neural Networks
by: Zanardini, Damiano, et al.
Published: (2024)
by: Zanardini, Damiano, et al.
Published: (2024)
Generative and Contrastive Graph Representation Learning
by: Chen, Jiali, et al.
Published: (2025)
by: Chen, Jiali, et al.
Published: (2025)
Comprehensive Metapath-based Heterogeneous Graph Transformer for Gene-Disease Association Prediction
by: Cui, Wentao, et al.
Published: (2025)
by: Cui, Wentao, et al.
Published: (2025)
Shattered Compositionality: Counterintuitive Learning Dynamics of Transformers for Arithmetic
by: Zhao, Xingyu, et al.
Published: (2026)
by: Zhao, Xingyu, et al.
Published: (2026)
Optimizing Robustness and Accuracy in Mixture of Experts: A Dual-Model Approach
by: Zhang, Xu, et al.
Published: (2025)
by: Zhang, Xu, et al.
Published: (2025)
HyDRA: Hybrid Dynamic Routing Architecture for Heterogeneous LLM Pools
by: Garg, Aashna, et al.
Published: (2026)
by: Garg, Aashna, et al.
Published: (2026)
CircuitProbe: Predicting Reasoning Circuits in Transformers via Stability Zone Detection
by: Panuganti, Rajkiran
Published: (2026)
by: Panuganti, Rajkiran
Published: (2026)
HEHRGNN: A Unified Embedding Model for Knowledge Graphs with Hyperedges and Hyper-Relational Edges
by: Rajagopalamenon, Rajesh, et al.
Published: (2026)
by: Rajagopalamenon, Rajesh, et al.
Published: (2026)
SMOSE: Sparse Mixture of Shallow Experts for Interpretable Reinforcement Learning in Continuous Control Tasks
by: Vincze, Mátyás, et al.
Published: (2024)
by: Vincze, Mátyás, et al.
Published: (2024)
One-vs.-One Mitigation of Intersectional Bias: A General Method to Extend Fairness-Aware Binary Classification
by: Kobayashi, Kenji, et al.
Published: (2020)
by: Kobayashi, Kenji, et al.
Published: (2020)
Hybrid-AIRL: Enhancing Inverse Reinforcement Learning with Supervised Expert Guidance
by: Silue, Bram, et al.
Published: (2025)
by: Silue, Bram, et al.
Published: (2025)
Entropy Causal Graphs for Multivariate Time Series Anomaly Detection
by: Febrinanto, Falih Gozi, et al.
Published: (2023)
by: Febrinanto, Falih Gozi, et al.
Published: (2023)
EEG Sleep Stage Classification with Continuous Wavelet Transform and Deep Learning
by: Gashti, Mehdi Zekriyapanah, et al.
Published: (2025)
by: Gashti, Mehdi Zekriyapanah, et al.
Published: (2025)
Fusion-Based Neural Generalization for Predicting Temperature Fields in Industrial PET Preform Heating
by: Alsheikh, Ahmad, et al.
Published: (2025)
by: Alsheikh, Ahmad, et al.
Published: (2025)
Towards Resource-Efficient Multimodal Intelligence: Learned Routing among Specialized Expert Models
by: Saini, Mayank, et al.
Published: (2025)
by: Saini, Mayank, et al.
Published: (2025)
Forget Attention: Importance-Aware Attention Is All You Need
by: Shin, Soohyeong, et al.
Published: (2026)
by: Shin, Soohyeong, et al.
Published: (2026)
Unsolvability Ceiling in Multi-LLM Routing: An Empirical Study of Evaluation Artifacts
by: Garg, Saloni, et al.
Published: (2026)
by: Garg, Saloni, et al.
Published: (2026)
TACO: Tackling Over-correction in Federated Learning with Tailored Adaptive Correction
by: Liu, Weijie, et al.
Published: (2025)
by: Liu, Weijie, et al.
Published: (2025)
DataRater: Meta-Learned Dataset Curation
by: Calian, Dan A., et al.
Published: (2025)
by: Calian, Dan A., et al.
Published: (2025)
The Lattice Geometry of Neural Network Quantization -- A Short Equivalence Proof of GPTQ and Babai's Algorithm
by: Birnick, Johann
Published: (2025)
by: Birnick, Johann
Published: (2025)
Unlearning's Blind Spots: Over-Unlearning and Prototypical Relearning Attack
by: Ha, SeungBum, et al.
Published: (2025)
by: Ha, SeungBum, et al.
Published: (2025)
Residual Reservoir Memory Networks
by: Pinna, Matteo, et al.
Published: (2025)
by: Pinna, Matteo, et al.
Published: (2025)
FreRA: A Frequency-Refined Augmentation for Contrastive Learning on Time Series Classification
by: Tian, Tian, et al.
Published: (2025)
by: Tian, Tian, et al.
Published: (2025)
Deep Residual Echo State Networks: exploring residual orthogonal connections in untrained Recurrent Neural Networks
by: Pinna, Matteo, et al.
Published: (2025)
by: Pinna, Matteo, et al.
Published: (2025)
1 bit is all we need: binary normalized neural networks
by: Cabral, Eduardo Lobo Lustoda, et al.
Published: (2025)
by: Cabral, Eduardo Lobo Lustoda, et al.
Published: (2025)
Multi-Task Reinforcement Learning with Language-Encoded Gated Policy Networks
by: Arora, Rushiv
Published: (2025)
by: Arora, Rushiv
Published: (2025)
Approximate Domain Unlearning for Vision-Language Models
by: Kawamura, Kodai, et al.
Published: (2025)
by: Kawamura, Kodai, et al.
Published: (2025)
Similar Items
-
Task-Conditioned Routing Signatures in Sparse Mixture-of-Experts Transformers
by: Avinash, Mynampati Sri Ranganadha
Published: (2026) -
ReflexGrad: Within-Episode Failure Recovery in LLM Agents via Progress-Gated Dual-Process Routing
by: Kadu, Ankush, et al.
Published: (2025) -
Deep Memory Search: A Metaheuristic Approach for Optimizing Heuristic Search
by: Hedar, Abdel-Rahman, et al.
Published: (2024) -
Hidden Failure Modes of Gradient Modification under Adam in Continual Learning, and Adaptive Decoupled Moment Routing as a Repair
by: Hu, Yuelin, et al.
Published: (2026) -
Not All Transitions Matter: Evidence from PPO
by: Basnet, Ajhesh
Published: (2026)