Unveiling Language Routing Isolation in Multilingual MoE Models for Interpretable Subnetwork Adaptation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zheng, Kening, Huang, Wei-Chieh, Huo, Jiahao, Li, Zhonghao, Zou, Henry Peng, Yan, Yibo, Zou, Xin, Li, Jungang, Li, Junzhuo, Zhang, Hanrong, Hu, Xuming, Yu, Philip S. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dynamic Expert Specialization: Towards Catastrophic Forgetting-Free Multi-Domain MoE Adaptation
von: Li, Junzhuo, et al.
Veröffentlicht: (2025)
von: Li, Junzhuo, et al.
Veröffentlicht: (2025)
Reefknot: A Comprehensive Benchmark for Relation Hallucination Evaluation, Analysis and Mitigation in Multimodal Large Language Models
von: Zheng, Kening, et al.
Veröffentlicht: (2024)
von: Zheng, Kening, et al.
Veröffentlicht: (2024)
Deconstructing Pre-training: Knowledge Attribution Analysis in MoE and Dense Models
von: Wang, Bo, et al.
Veröffentlicht: (2026)
von: Wang, Bo, et al.
Veröffentlicht: (2026)
GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning
von: Zhang, Jianghangfan, et al.
Veröffentlicht: (2025)
von: Zhang, Jianghangfan, et al.
Veröffentlicht: (2025)
Pierce the Mists, Greet the Sky: Decipher Knowledge Overshadowing via Knowledge Circuit Analysis
von: Huang, Haoming, et al.
Veröffentlicht: (2025)
von: Huang, Haoming, et al.
Veröffentlicht: (2025)
Unlocking Multimodal Document Intelligence: From Current Triumphs to Future Frontiers of Visual Document Retrieval
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
Unveiling Instruction-Specific Neurons & Experts: An Analytical Framework for LLM's Instruction-Following Capabilities
von: Zhang, Junyan, et al.
Veröffentlicht: (2025)
von: Zhang, Junyan, et al.
Veröffentlicht: (2025)
Capturing Nuanced Preferences: Preference-Aligned Distillation for Small Language Models
von: Gu, Yanggan, et al.
Veröffentlicht: (2025)
von: Gu, Yanggan, et al.
Veröffentlicht: (2025)
SAMark: A Self-Anchored Text Watermarking with Paragraph-Level Paraphrase Robustness
von: Huo, Jiahao, et al.
Veröffentlicht: (2026)
von: Huo, Jiahao, et al.
Veröffentlicht: (2026)
CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding
von: Huo, Jiahao, et al.
Veröffentlicht: (2026)
von: Huo, Jiahao, et al.
Veröffentlicht: (2026)
MMNeuron: Discovering Neuron-Level Domain-Specific Interpretation in Multimodal Large Language Model
von: Huo, Jiahao, et al.
Veröffentlicht: (2024)
von: Huo, Jiahao, et al.
Veröffentlicht: (2024)
MMUnlearner: Reformulating Multimodal Machine Unlearning in the Era of Multimodal Large Language Models
von: Huo, Jiahao, et al.
Veröffentlicht: (2025)
von: Huo, Jiahao, et al.
Veröffentlicht: (2025)
Refiner: Restructure Retrieval Content Efficiently to Advance Question-Answering Capabilities
von: Li, Zhonghao, et al.
Veröffentlicht: (2024)
von: Li, Zhonghao, et al.
Veröffentlicht: (2024)
CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification
von: Zhang, Hanrong, et al.
Veröffentlicht: (2026)
von: Zhang, Hanrong, et al.
Veröffentlicht: (2026)
Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing Guidance
von: Wei, Yujie, et al.
Veröffentlicht: (2025)
von: Wei, Yujie, et al.
Veröffentlicht: (2025)
Visual Late Chunking: An Empirical Study of Contextual Chunking for Efficient Visual Document Retrieval
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
Sculpting the Vector Space: Towards Efficient Multi-Vector Visual Document Retrieval via Prune-then-Merge Framework
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
Mix-MoE: Improving Multilingual Machine Translation of Large Language Models through Mixed MoEs
von: Li, Bo, et al.
Veröffentlicht: (2026)
von: Li, Bo, et al.
Veröffentlicht: (2026)
LoTA-QAF: Lossless Ternary Adaptation for Quantization-Aware Fine-Tuning
von: Chen, Junyu, et al.
Veröffentlicht: (2025)
von: Chen, Junyu, et al.
Veröffentlicht: (2025)
Accelerating Distributed MoE Training and Inference with Lina
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
Internal Chain-of-Thought: Empirical Evidence for Layer-wise Subtask Scheduling in LLMs
von: Yang, Zhipeng, et al.
Veröffentlicht: (2025)
von: Yang, Zhipeng, et al.
Veröffentlicht: (2025)
BLR-MoE: Boosted Language-Routing Mixture of Experts for Domain-Robust Multilingual E2E ASR
von: Ma, Guodong, et al.
Veröffentlicht: (2025)
von: Ma, Guodong, et al.
Veröffentlicht: (2025)
GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference
von: Han, Yu, et al.
Veröffentlicht: (2025)
von: Han, Yu, et al.
Veröffentlicht: (2025)
MathAgent: Leveraging a Mixture-of-Math-Agent Framework for Real-World Multimodal Mathematical Error Detection
von: Yan, Yibo, et al.
Veröffentlicht: (2025)
von: Yan, Yibo, et al.
Veröffentlicht: (2025)
Beyond the Grid: Layout-Informed Multi-Vector Retrieval with Parsed Visual Document Representations
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
MoE-LPR: Multilingual Extension of Large Language Models through Mixture-of-Experts with Language Priors Routing
von: Zhou, Hao, et al.
Veröffentlicht: (2024)
von: Zhou, Hao, et al.
Veröffentlicht: (2024)
ARC-STAR: Auditable Post-Hoc Correction for PDE Foundation Models
von: Li, Chengze, et al.
Veröffentlicht: (2026)
von: Li, Chengze, et al.
Veröffentlicht: (2026)
Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning
von: Qin, Jialong, et al.
Veröffentlicht: (2025)
von: Qin, Jialong, et al.
Veröffentlicht: (2025)
On-the-fly Routing for Zero-shot MoE Speaker Adaptation of Speech Foundation Models for Dysarthric Speech Recognition
von: HU, Shujie, et al.
Veröffentlicht: (2025)
von: HU, Shujie, et al.
Veröffentlicht: (2025)
DynaMo: Runtime Switchable Quantization for MoE with Cross-Dataset Adaptation
von: Zheng, Zihao, et al.
Veröffentlicht: (2025)
von: Zheng, Zihao, et al.
Veröffentlicht: (2025)
MoE-Sieve: Routing-Guided LoRA for Efficient MoE Fine-Tuning
von: Manzoni, Andrea
Veröffentlicht: (2026)
von: Manzoni, Andrea
Veröffentlicht: (2026)
BEAM: Binary Expert Activation Masking for Dynamic Routing in MoE
von: Wu, Juntong, et al.
Veröffentlicht: (2026)
von: Wu, Juntong, et al.
Veröffentlicht: (2026)
Exploring Response Uncertainty in MLLMs: An Empirical Evaluation under Misleading Scenarios
von: Dang, Yunkai, et al.
Veröffentlicht: (2024)
von: Dang, Yunkai, et al.
Veröffentlicht: (2024)
Expert-Token Resonance MoE: Bidirectional Routing with Efficiency Affinity-Driven Active Selection
von: Li, Jing, et al.
Veröffentlicht: (2024)
von: Li, Jing, et al.
Veröffentlicht: (2024)
Awakening Dormant Experts:Counterfactual Routing to Mitigate MoE Hallucinations
von: Hu, Wentao, et al.
Veröffentlicht: (2026)
von: Hu, Wentao, et al.
Veröffentlicht: (2026)
I2MoE: Interpretable Multimodal Interaction-aware Mixture-of-Experts
von: Xin, Jiayi, et al.
Veröffentlicht: (2025)
von: Xin, Jiayi, et al.
Veröffentlicht: (2025)
BIG-MoE: Bypass Isolated Gating MoE for Generalized Multimodal Face Anti-Spoofing
von: Ma, Yingjie, et al.
Veröffentlicht: (2024)
von: Ma, Yingjie, et al.
Veröffentlicht: (2024)
MiM-DiT: MoE in MoE with Diffusion Transformers for All-in-One Image Restoration
von: Kong, Lingshun, et al.
Veröffentlicht: (2026)
von: Kong, Lingshun, et al.
Veröffentlicht: (2026)
Leave It to the Experts: Detecting Knowledge Distillation via MoE Expert Signatures
von: Li, Pingzhi, et al.
Veröffentlicht: (2025)
von: Li, Pingzhi, et al.
Veröffentlicht: (2025)
Temporal Gains, Spatial Costs: Revisiting Video Fine-Tuning in Multimodal Large Language Models
von: Zhang, Linghao, et al.
Veröffentlicht: (2026)
von: Zhang, Linghao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Dynamic Expert Specialization: Towards Catastrophic Forgetting-Free Multi-Domain MoE Adaptation
von: Li, Junzhuo, et al.
Veröffentlicht: (2025) -
Reefknot: A Comprehensive Benchmark for Relation Hallucination Evaluation, Analysis and Mitigation in Multimodal Large Language Models
von: Zheng, Kening, et al.
Veröffentlicht: (2024) -
Deconstructing Pre-training: Knowledge Attribution Analysis in MoE and Dense Models
von: Wang, Bo, et al.
Veröffentlicht: (2026) -
GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning
von: Zhang, Jianghangfan, et al.
Veröffentlicht: (2025) -
Pierce the Mists, Greet the Sky: Decipher Knowledge Overshadowing via Knowledge Circuit Analysis
von: Huang, Haoming, et al.
Veröffentlicht: (2025)