Layer-wise MoE Routing Locality under Shared-Prefix Code Generation: Token-Identity Decomposition and Compile-Equivalent Fork Redundancy
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hayashi, Shun-ichiro, Mukunoki, Daichi, Hoshino, Tetsuya, Katagiri, Takahiro |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Performance Evaluation of General Purpose Large Language Models for Basic Linear Algebra Subprograms Code Generation
von: Mukunoki, Daichi, et al.
Veröffentlicht: (2025)
von: Mukunoki, Daichi, et al.
Veröffentlicht: (2025)
VibeCodeHPC: An Agent-Based Iterative Prompting Auto-Tuner for HPC Code Generation Using LLMs
von: Hayashi, Shun-ichiro, et al.
Veröffentlicht: (2025)
von: Hayashi, Shun-ichiro, et al.
Veröffentlicht: (2025)
Improving HPC Code Generation Capability of LLMs via Online Reinforcement Learning with Real-Machine Benchmark Rewards
von: Mikasa, Ryo, et al.
Veröffentlicht: (2026)
von: Mikasa, Ryo, et al.
Veröffentlicht: (2026)
3Dify: a Framework for Procedural 3D-CG Generation Assisted by LLMs Using MCP and RAG
von: Hayashi, Shun-ichiro, et al.
Veröffentlicht: (2025)
von: Hayashi, Shun-ichiro, et al.
Veröffentlicht: (2025)
Towards Generalized Parameter Tuning in Coherent Ising Machines: A Portfolio-Based Approach
von: Hanyu, Tatsuro, et al.
Veröffentlicht: (2025)
von: Hanyu, Tatsuro, et al.
Veröffentlicht: (2025)
Learning-Augmented Performance Model for Tensor Product Factorization in High-Order FEM
von: Ren, Xuanzhengbo, et al.
Veröffentlicht: (2026)
von: Ren, Xuanzhengbo, et al.
Veröffentlicht: (2026)
DGEMM without FP64 Arithmetic - Using FP64 Emulation and FP8 Tensor Cores with Ozaki Scheme
von: Mukunoki, Daichi
Veröffentlicht: (2025)
von: Mukunoki, Daichi
Veröffentlicht: (2025)
Sparse Iterative Solvers Using High-Precision Arithmetic with Quasi Multi-Word Algorithms
von: Mukunoki, Daichi, et al.
Veröffentlicht: (2025)
von: Mukunoki, Daichi, et al.
Veröffentlicht: (2025)
MoE-Prefill: Zero Redundancy Overheads in MoE Prefill Serving
von: Su, Zhaoyuan, et al.
Veröffentlicht: (2026)
von: Su, Zhaoyuan, et al.
Veröffentlicht: (2026)
GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference
von: Han, Yu, et al.
Veröffentlicht: (2025)
von: Han, Yu, et al.
Veröffentlicht: (2025)
Adaptation of XAI to Auto-tuning for Numerical Libraries
von: Aoki, Shota, et al.
Veröffentlicht: (2024)
von: Aoki, Shota, et al.
Veröffentlicht: (2024)
Performance Evaluation of CMOS Annealing with Support Vector Machine
von: Fukuhara, Ryoga, et al.
Veröffentlicht: (2024)
von: Fukuhara, Ryoga, et al.
Veröffentlicht: (2024)
From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill
von: Lee, Gunjun, et al.
Veröffentlicht: (2025)
von: Lee, Gunjun, et al.
Veröffentlicht: (2025)
MoNE: Replacing Redundant Experts with Lightweight Novices for Structured Pruning of MoE
von: Zhang, Geng, et al.
Veröffentlicht: (2025)
von: Zhang, Geng, et al.
Veröffentlicht: (2025)
Expert-Token Resonance MoE: Bidirectional Routing with Efficiency Affinity-Driven Active Selection
von: Li, Jing, et al.
Veröffentlicht: (2024)
von: Li, Jing, et al.
Veröffentlicht: (2024)
MoE-Sieve: Routing-Guided LoRA for Efficient MoE Fine-Tuning
von: Manzoni, Andrea
Veröffentlicht: (2026)
von: Manzoni, Andrea
Veröffentlicht: (2026)
ppOpen-AT: A Directive-base Auto-tuning Language
von: Katagiri, Takahiro
Veröffentlicht: (2024)
von: Katagiri, Takahiro
Veröffentlicht: (2024)
Local Equivalence Problem in Hidden Markov Model
von: Hayashi, Masahito
Veröffentlicht: (2018)
von: Hayashi, Masahito
Veröffentlicht: (2018)
Stable-MoE: Lyapunov-based Token Routing for Distributed Mixture-of-Experts Training over Edge Networks
von: Shi, Long, et al.
Veröffentlicht: (2025)
von: Shi, Long, et al.
Veröffentlicht: (2025)
HiFi-MambaV2: Hierarchical Shared-Routed MoE for High-Fidelity MRI Reconstruction
von: Fang, Pengcheng, et al.
Veröffentlicht: (2025)
von: Fang, Pengcheng, et al.
Veröffentlicht: (2025)
Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing Guidance
von: Wei, Yujie, et al.
Veröffentlicht: (2025)
von: Wei, Yujie, et al.
Veröffentlicht: (2025)
R^2MoE: Redundancy-Removal Mixture of Experts for Lifelong Concept Learning
von: Guo, Xiaohan, et al.
Veröffentlicht: (2025)
von: Guo, Xiaohan, et al.
Veröffentlicht: (2025)
D$^{2}$MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
von: Wang, Haodong, et al.
Veröffentlicht: (2025)
von: Wang, Haodong, et al.
Veröffentlicht: (2025)
Dynamic Language Group-Based MoE: Enhancing Code-Switching Speech Recognition with Hierarchical Routing
von: Huang, Hukai, et al.
Veröffentlicht: (2024)
von: Huang, Hukai, et al.
Veröffentlicht: (2024)
Polysemantic Experts, Monosemantic Paths: Routing as Control in MoEs
von: Ye, Charles, et al.
Veröffentlicht: (2026)
von: Ye, Charles, et al.
Veröffentlicht: (2026)
LSH-MoE: Communication-efficient MoE Training via Locality-Sensitive Hashing
von: Nie, Xiaonan, et al.
Veröffentlicht: (2024)
von: Nie, Xiaonan, et al.
Veröffentlicht: (2024)
SD-MoE: Spectral Decomposition for Effective Expert Specialization
von: Huang, Ruijun, et al.
Veröffentlicht: (2026)
von: Huang, Ruijun, et al.
Veröffentlicht: (2026)
MoE-nD: Per-Layer Mixture-of-Experts Routing for Multi-Axis KV Cache Compression
von: Sun, Libo, et al.
Veröffentlicht: (2026)
von: Sun, Libo, et al.
Veröffentlicht: (2026)
MoE Routing Testbed: Studying Expert Specialization and Routing Behavior at Small Scale
von: Falke, Tobias, et al.
Veröffentlicht: (2026)
von: Falke, Tobias, et al.
Veröffentlicht: (2026)
Theories of Frege structure equivalent to Feferman's system $\mathsf{T}_0$
von: Hayashi, Daichi
Veröffentlicht: (2024)
von: Hayashi, Daichi
Veröffentlicht: (2024)
Guiding the Experts: Semantic Priors for Efficient and Focused MoE Routing
von: Min, Chengxi, et al.
Veröffentlicht: (2025)
von: Min, Chengxi, et al.
Veröffentlicht: (2025)
BEAM: Binary Expert Activation Masking for Dynamic Routing in MoE
von: Wu, Juntong, et al.
Veröffentlicht: (2026)
von: Wu, Juntong, et al.
Veröffentlicht: (2026)
Statistic-Augmented, Decoupled MoE Routing and Aggregating in Autonomous Driving
von: Kou, Wei-Bin, et al.
Veröffentlicht: (2025)
von: Kou, Wei-Bin, et al.
Veröffentlicht: (2025)
Grouter: Decoupling Routing from Representation for Accelerated MoE Training
von: Xu, Yuqi, et al.
Veröffentlicht: (2026)
von: Xu, Yuqi, et al.
Veröffentlicht: (2026)
Awakening Dormant Experts:Counterfactual Routing to Mitigate MoE Hallucinations
von: Hu, Wentao, et al.
Veröffentlicht: (2026)
von: Hu, Wentao, et al.
Veröffentlicht: (2026)
Ada-K Routing: Boosting the Efficiency of MoE-based LLMs
von: Yue, Tongtian, et al.
Veröffentlicht: (2024)
von: Yue, Tongtian, et al.
Veröffentlicht: (2024)
MoE-DP: An MoE-Enhanced Diffusion Policy for Robust Long-Horizon Robotic Manipulation with Skill Decomposition and Failure Recovery
von: Cheng, Baiye, et al.
Veröffentlicht: (2025)
von: Cheng, Baiye, et al.
Veröffentlicht: (2025)
Prefix Grouper: Efficient GRPO Training through Shared-Prefix Forward
von: Liu, Zikang, et al.
Veröffentlicht: (2025)
von: Liu, Zikang, et al.
Veröffentlicht: (2025)
Mixture of Tokens: Continuous MoE through Cross-Example Aggregation
von: Antoniak, Szymon, et al.
Veröffentlicht: (2023)
von: Antoniak, Szymon, et al.
Veröffentlicht: (2023)
XShare: Collaborative in-Batch Expert Sharing for Faster MoE Inference
von: Vankov, Daniil, et al.
Veröffentlicht: (2026)
von: Vankov, Daniil, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Performance Evaluation of General Purpose Large Language Models for Basic Linear Algebra Subprograms Code Generation
von: Mukunoki, Daichi, et al.
Veröffentlicht: (2025) -
VibeCodeHPC: An Agent-Based Iterative Prompting Auto-Tuner for HPC Code Generation Using LLMs
von: Hayashi, Shun-ichiro, et al.
Veröffentlicht: (2025) -
Improving HPC Code Generation Capability of LLMs via Online Reinforcement Learning with Real-Machine Benchmark Rewards
von: Mikasa, Ryo, et al.
Veröffentlicht: (2026) -
3Dify: a Framework for Procedural 3D-CG Generation Assisted by LLMs Using MCP and RAG
von: Hayashi, Shun-ichiro, et al.
Veröffentlicht: (2025) -
Towards Generalized Parameter Tuning in Coherent Ising Machines: A Portfolio-Based Approach
von: Hanyu, Tatsuro, et al.
Veröffentlicht: (2025)