Cosine-Similarity Routing with Semantic Anchors for Interpretable Mixture-of-Experts Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Ternovtsii, Ivan, Bilak, Yurii |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Equifinality in Mixture of Experts: Routing Topology Does Not Determine Language Modeling Quality
by: Ternovtsii, Ivan, et al.
Published: (2026)
by: Ternovtsii, Ivan, et al.
Published: (2026)
Geometric Routing Enables Causal Expert Control in Mixture of Experts
by: Ternovtsii, Ivan, et al.
Published: (2026)
by: Ternovtsii, Ivan, et al.
Published: (2026)
Mixture of Layers with Hybrid Attention
by: Ternovtsii, Ivan, et al.
Published: (2026)
by: Ternovtsii, Ivan, et al.
Published: (2026)
Probing Semantic Routing in Large Mixture-of-Expert Models
by: Olson, Matthew Lyle, et al.
Published: (2025)
by: Olson, Matthew Lyle, et al.
Published: (2025)
Multilingual Routing in Mixture-of-Experts
by: Bandarkar, Lucas, et al.
Published: (2025)
by: Bandarkar, Lucas, et al.
Published: (2025)
Routing-Free Mixture-of-Experts
by: Liu, Yilun, et al.
Published: (2026)
by: Liu, Yilun, et al.
Published: (2026)
Routing by Analogy: kNN-Augmented Expert Assignment for Mixture-of-Experts
by: Lyu, Boxuan, et al.
Published: (2026)
by: Lyu, Boxuan, et al.
Published: (2026)
Routing-Aligned Fine-Tuning for Multilingual Downstream Tasks in Mixture-of-Experts Models
by: Deng, Guanzhi, et al.
Published: (2026)
by: Deng, Guanzhi, et al.
Published: (2026)
The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert Level
by: Herbst, Jeremy, et al.
Published: (2026)
by: Herbst, Jeremy, et al.
Published: (2026)
MoxE: Mixture of xLSTM Experts with Entropy-Aware Routing for Efficient Language Modeling
by: Thiombiano, Abdoul Majid O., et al.
Published: (2025)
by: Thiombiano, Abdoul Majid O., et al.
Published: (2025)
Pre-Attention Expert Prediction and Prefetching for Mixture-of-Experts Large Language Models
by: Zhu, Shien, et al.
Published: (2025)
by: Zhu, Shien, et al.
Published: (2025)
LD-MoLE: Learnable Dynamic Routing for Mixture of LoRA Experts
by: Zhuang, Yuan, et al.
Published: (2025)
by: Zhuang, Yuan, et al.
Published: (2025)
Laying Anchors: Semantically Priming Numerals in Language Modeling
by: Sharma, Mandar, et al.
Published: (2024)
by: Sharma, Mandar, et al.
Published: (2024)
Group then Scale: Dynamic Mixture-of-Experts Multilingual Language Model
by: Li, Chong, et al.
Published: (2025)
by: Li, Chong, et al.
Published: (2025)
Diversifying the Mixture-of-Experts Representation for Language Models with Orthogonal Optimizer
by: Liu, Boan, et al.
Published: (2023)
by: Liu, Boan, et al.
Published: (2023)
Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts
by: Xu, Haolei, et al.
Published: (2026)
by: Xu, Haolei, et al.
Published: (2026)
GigaChat Family: Efficient Russian Language Modeling Through Mixture of Experts Architecture
by: GigaChat team, et al.
Published: (2025)
by: GigaChat team, et al.
Published: (2025)
Parameter-Efficient Routed Fine-Tuning: Mixture-of-Experts Demands Mixture of Adaptation Modules
by: Liu, Yilun, et al.
Published: (2025)
by: Liu, Yilun, et al.
Published: (2025)
Information Flow Routes: Automatically Interpreting Language Models at Scale
by: Ferrando, Javier, et al.
Published: (2024)
by: Ferrando, Javier, et al.
Published: (2024)
On the Spatial Structure of Mixture-of-Experts in Transformers
by: Bershatsky, Daniel, et al.
Published: (2025)
by: Bershatsky, Daniel, et al.
Published: (2025)
SciDFM: A Large Language Model with Mixture-of-Experts for Science
by: Sun, Liangtai, et al.
Published: (2024)
by: Sun, Liangtai, et al.
Published: (2024)
Mixture of Heterogeneous Grouped Experts for Language Modeling
by: Ma, Zhicheng, et al.
Published: (2026)
by: Ma, Zhicheng, et al.
Published: (2026)
Upcycling Large Language Models into Mixture of Experts
by: He, Ethan, et al.
Published: (2024)
by: He, Ethan, et al.
Published: (2024)
SAC-Opt: Semantic Anchors for Iterative Correction in Optimization Modeling
by: Zhang, Yansen, et al.
Published: (2025)
by: Zhang, Yansen, et al.
Published: (2025)
Anchor-based Large Language Models
by: Pang, Jianhui, et al.
Published: (2024)
by: Pang, Jianhui, et al.
Published: (2024)
Flexible and Effective Mixing of Large Language Models into a Mixture of Domain Experts
by: Lee, Rhui Dih, et al.
Published: (2024)
by: Lee, Rhui Dih, et al.
Published: (2024)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
by: Kim, Junhyuck, et al.
Published: (2026)
by: Kim, Junhyuck, et al.
Published: (2026)
OLMoE: Open Mixture-of-Experts Language Models
by: Muennighoff, Niklas, et al.
Published: (2024)
by: Muennighoff, Niklas, et al.
Published: (2024)
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
by: Lu, Xudong, et al.
Published: (2024)
by: Lu, Xudong, et al.
Published: (2024)
Expert Threshold Routing for Autoregressive Language Modeling with Dynamic Computation Allocation and Load Balancing
by: Sun, Hanchi, et al.
Published: (2026)
by: Sun, Hanchi, et al.
Published: (2026)
SAMoRA: Semantic-Aware Mixture of LoRA Experts for Task-Adaptive Learning
by: Shi, Boyan, et al.
Published: (2026)
by: Shi, Boyan, et al.
Published: (2026)
Marco-MoE: Open Multilingual Mixture-of-Expert Language Models with Efficient Upcycling
by: Jiang, Fan, et al.
Published: (2026)
by: Jiang, Fan, et al.
Published: (2026)
DynMoLE: Boosting Mixture of LoRA Experts Fine-Tuning with a Hybrid Routing Mechanism
by: Li, Dengchun, et al.
Published: (2025)
by: Li, Dengchun, et al.
Published: (2025)
Every Expert Matters: Towards Effective Knowledge Distillation for Mixture-of-Experts Language Models
by: Kim, Gyeongman, et al.
Published: (2025)
by: Kim, Gyeongman, et al.
Published: (2025)
Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
by: Nakamura, Taishi, et al.
Published: (2025)
by: Nakamura, Taishi, et al.
Published: (2025)
SparseDoctor: Towards Efficient Chat Doctor with Mixture of Experts Enhanced Large Language Models
by: Zhang, Jianbin, et al.
Published: (2025)
by: Zhang, Jianbin, et al.
Published: (2025)
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
by: DeepSeek-AI, et al.
Published: (2024)
by: DeepSeek-AI, et al.
Published: (2024)
Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Language Models
by: Wei, Tianwen, et al.
Published: (2024)
by: Wei, Tianwen, et al.
Published: (2024)
Unveiling Language Routing Isolation in Multilingual MoE Models for Interpretable Subnetwork Adaptation
by: Zheng, Kening, et al.
Published: (2026)
by: Zheng, Kening, et al.
Published: (2026)
Skill-Based Mixture-of-Experts: Adaptive Routing for Heterogeneous Reasoning via Inferred Skills
by: Chen, Justin Chih-Yao, et al.
Published: (2025)
by: Chen, Justin Chih-Yao, et al.
Published: (2025)
Similar Items
-
Equifinality in Mixture of Experts: Routing Topology Does Not Determine Language Modeling Quality
by: Ternovtsii, Ivan, et al.
Published: (2026) -
Geometric Routing Enables Causal Expert Control in Mixture of Experts
by: Ternovtsii, Ivan, et al.
Published: (2026) -
Mixture of Layers with Hybrid Attention
by: Ternovtsii, Ivan, et al.
Published: (2026) -
Probing Semantic Routing in Large Mixture-of-Expert Models
by: Olson, Matthew Lyle, et al.
Published: (2025) -
Multilingual Routing in Mixture-of-Experts
by: Bandarkar, Lucas, et al.
Published: (2025)