Decoding Knowledge Attribution in Mixture-of-Experts: A Framework of Basic-Refinement Collaboration and Efficiency Analysis
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Junzhuo, Wang, Bo, Zhou, Xiuze, Jiang, Peijie, Liu, Jia, Hu, Xuming |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Dynamic Expert Specialization: Towards Catastrophic Forgetting-Free Multi-Domain MoE Adaptation
por: Li, Junzhuo, et al.
Publicado: (2025)
por: Li, Junzhuo, et al.
Publicado: (2025)
Optimal Expert-Attention Allocation in Mixture-of-Experts: A Scalable Law for Dynamic Model Design
por: Li, Junzhuo, et al.
Publicado: (2026)
por: Li, Junzhuo, et al.
Publicado: (2026)
Deconstructing Pre-training: Knowledge Attribution Analysis in MoE and Dense Models
por: Wang, Bo, et al.
Publicado: (2026)
por: Wang, Bo, et al.
Publicado: (2026)
Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression
por: Liu, Xiang, et al.
Publicado: (2025)
por: Liu, Xiang, et al.
Publicado: (2025)
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
por: Lu, Xudong, et al.
Publicado: (2024)
por: Lu, Xudong, et al.
Publicado: (2024)
SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?
por: Zhuang, Haomin, et al.
Publicado: (2024)
por: Zhuang, Haomin, et al.
Publicado: (2024)
Multilingual Routing in Mixture-of-Experts
por: Bandarkar, Lucas, et al.
Publicado: (2025)
por: Bandarkar, Lucas, et al.
Publicado: (2025)
MoECollab: Democratizing LLM Development Through Collaborative Mixture of Experts
por: Harshit
Publicado: (2025)
por: Harshit
Publicado: (2025)
Every Expert Matters: Towards Effective Knowledge Distillation for Mixture-of-Experts Language Models
por: Kim, Gyeongman, et al.
Publicado: (2025)
por: Kim, Gyeongman, et al.
Publicado: (2025)
LongGenBench: Long-context Generation Benchmark
por: Liu, Xiang, et al.
Publicado: (2024)
por: Liu, Xiang, et al.
Publicado: (2024)
Exploring Response Uncertainty in MLLMs: An Empirical Evaluation under Misleading Scenarios
por: Dang, Yunkai, et al.
Publicado: (2024)
por: Dang, Yunkai, et al.
Publicado: (2024)
Routing-Free Mixture-of-Experts
por: Liu, Yilun, et al.
Publicado: (2026)
por: Liu, Yilun, et al.
Publicado: (2026)
Knowledge Localization in Mixture-of-Experts LLMs Using Cross-Lingual Inconsistency
por: Bandarkar, Lucas, et al.
Publicado: (2026)
por: Bandarkar, Lucas, et al.
Publicado: (2026)
Evolutionary Guided Decoding: Iterative Value Refinement for LLMs
por: Liu, Zhenhua, et al.
Publicado: (2025)
por: Liu, Zhenhua, et al.
Publicado: (2025)
FactorLLM: Factorizing Knowledge via Mixture of Experts for Large Language Models
por: Zhao, Zhongyu, et al.
Publicado: (2024)
por: Zhao, Zhongyu, et al.
Publicado: (2024)
COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation
por: Zhou, Tianyi, et al.
Publicado: (2026)
por: Zhou, Tianyi, et al.
Publicado: (2026)
A Unified Virtual Mixture-of-Experts Framework:Enhanced Inference and Hallucination Mitigation in Single-Model System
por: Liu, Mingyan
Publicado: (2025)
por: Liu, Mingyan
Publicado: (2025)
Mixture of Heterogeneous Grouped Experts for Language Modeling
por: Ma, Zhicheng, et al.
Publicado: (2026)
por: Ma, Zhicheng, et al.
Publicado: (2026)
Multi-Head Mixture-of-Experts
por: Wu, Xun, et al.
Publicado: (2024)
por: Wu, Xun, et al.
Publicado: (2024)
Mixture of Attentions For Speculative Decoding
por: Zimmer, Matthieu, et al.
Publicado: (2024)
por: Zimmer, Matthieu, et al.
Publicado: (2024)
Upcycling Large Language Models into Mixture of Experts
por: He, Ethan, et al.
Publicado: (2024)
por: He, Ethan, et al.
Publicado: (2024)
Capturing Nuanced Preferences: Preference-Aligned Distillation for Small Language Models
por: Gu, Yanggan, et al.
Publicado: (2025)
por: Gu, Yanggan, et al.
Publicado: (2025)
Parameter-Efficient Routed Fine-Tuning: Mixture-of-Experts Demands Mixture of Adaptation Modules
por: Liu, Yilun, et al.
Publicado: (2025)
por: Liu, Yilun, et al.
Publicado: (2025)
Towards a Comprehensive Scaling Law of Mixture-of-Experts
por: Zhao, Guoliang, et al.
Publicado: (2025)
por: Zhao, Guoliang, et al.
Publicado: (2025)
FlyLoRA: Boosting Task Decoupling and Parameter Efficiency via Implicit Rank-Wise Mixture-of-Experts
por: Zou, Heming, et al.
Publicado: (2025)
por: Zou, Heming, et al.
Publicado: (2025)
On the Spatial Structure of Mixture-of-Experts in Transformers
por: Bershatsky, Daniel, et al.
Publicado: (2025)
por: Bershatsky, Daniel, et al.
Publicado: (2025)
MobileMoE: Scaling On-Device Mixture of Experts
por: Chen, Yanbei, et al.
Publicado: (2026)
por: Chen, Yanbei, et al.
Publicado: (2026)
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
por: Liu, Baihui, et al.
Publicado: (2026)
por: Liu, Baihui, et al.
Publicado: (2026)
Scaling Laws for Fine-Grained Mixture of Experts
por: Krajewski, Jakub, et al.
Publicado: (2024)
por: Krajewski, Jakub, et al.
Publicado: (2024)
MoIN: Mixture of Introvert Experts to Upcycle an LLM
por: Tejankar, Ajinkya, et al.
Publicado: (2024)
por: Tejankar, Ajinkya, et al.
Publicado: (2024)
Dynamic Experts Search: Enhancing Reasoning in Mixture-of-Experts LLMs at Test Time
por: Han, Yixuan, et al.
Publicado: (2025)
por: Han, Yixuan, et al.
Publicado: (2025)
Fast Large Language Model Collaborative Decoding via Speculation
por: Fu, Jiale, et al.
Publicado: (2025)
por: Fu, Jiale, et al.
Publicado: (2025)
MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapper
por: Zeng, Runjia, et al.
Publicado: (2025)
por: Zeng, Runjia, et al.
Publicado: (2025)
Mixture-of-Experts as Soft Clustering: A Dual Jacobian-PCA Spectral Geometry Perspective
por: Liu, Feilong
Publicado: (2026)
por: Liu, Feilong
Publicado: (2026)
Reflection-Window Decoding: Text Generation with Selective Refinement
por: Tang, Zeyu, et al.
Publicado: (2025)
por: Tang, Zeyu, et al.
Publicado: (2025)
$μ$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts
por: Koike-Akino, Toshiaki, et al.
Publicado: (2025)
por: Koike-Akino, Toshiaki, et al.
Publicado: (2025)
Mixtures of SubExperts for Large Language Continual Learning
por: Kang, Haeyong
Publicado: (2025)
por: Kang, Haeyong
Publicado: (2025)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
por: Kim, Junhyuck, et al.
Publicado: (2026)
por: Kim, Junhyuck, et al.
Publicado: (2026)
Probing Semantic Routing in Large Mixture-of-Expert Models
por: Olson, Matthew Lyle, et al.
Publicado: (2025)
por: Olson, Matthew Lyle, et al.
Publicado: (2025)
OLMoE: Open Mixture-of-Experts Language Models
por: Muennighoff, Niklas, et al.
Publicado: (2024)
por: Muennighoff, Niklas, et al.
Publicado: (2024)
Ejemplares similares
-
Dynamic Expert Specialization: Towards Catastrophic Forgetting-Free Multi-Domain MoE Adaptation
por: Li, Junzhuo, et al.
Publicado: (2025) -
Optimal Expert-Attention Allocation in Mixture-of-Experts: A Scalable Law for Dynamic Model Design
por: Li, Junzhuo, et al.
Publicado: (2026) -
Deconstructing Pre-training: Knowledge Attribution Analysis in MoE and Dense Models
por: Wang, Bo, et al.
Publicado: (2026) -
Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression
por: Liu, Xiang, et al.
Publicado: (2025) -
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
por: Lu, Xudong, et al.
Publicado: (2024)