Part-Of-Speech Sensitivity of Routers in Mixture of Experts Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Antoine, Elie, Béchet, Frédéric, Langlais, Philippe |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A linguistically-motivated evaluation methodology for unraveling model's abilities in reading comprehension tasks
di: Antoine, Elie, et al.
Pubblicazione: (2025)
di: Antoine, Elie, et al.
Pubblicazione: (2025)
Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling
di: Ran, Junfeng, et al.
Pubblicazione: (2025)
di: Ran, Junfeng, et al.
Pubblicazione: (2025)
Layerwise Recurrent Router for Mixture-of-Experts
di: Qiu, Zihan, et al.
Pubblicazione: (2024)
di: Qiu, Zihan, et al.
Pubblicazione: (2024)
On Evaluation Protocols for Data Augmentation in a Limited Data Scenario
di: Piedboeuf, Frédéric, et al.
Pubblicazione: (2024)
di: Piedboeuf, Frédéric, et al.
Pubblicazione: (2024)
Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss
di: Lv, Ang, et al.
Pubblicazione: (2025)
di: Lv, Ang, et al.
Pubblicazione: (2025)
Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts
di: Ahrac, Sagi, et al.
Pubblicazione: (2026)
di: Ahrac, Sagi, et al.
Pubblicazione: (2026)
Omni-Router: Sharing Routing Decisions in Sparse Mixture-of-Experts for Speech Recognition
di: Gu, Zijin, et al.
Pubblicazione: (2025)
di: Gu, Zijin, et al.
Pubblicazione: (2025)
Mixture of Routers
di: Zhang, Jia-Chen, et al.
Pubblicazione: (2025)
di: Zhang, Jia-Chen, et al.
Pubblicazione: (2025)
$\textit{BenchIE}^{FL}$ : A Manually Re-Annotated Fact-Based Open Information Extraction Benchmark
di: Lamarche, Fabrice, et al.
Pubblicazione: (2024)
di: Lamarche, Fabrice, et al.
Pubblicazione: (2024)
Yuan 2.0-M32: Mixture of Experts with Attention Router
di: Wu, Shaohua, et al.
Pubblicazione: (2024)
di: Wu, Shaohua, et al.
Pubblicazione: (2024)
Decomposing Retrieval Failures in RAG for Long-Document Financial Question Answering
di: Kobeissi, Amine, et al.
Pubblicazione: (2026)
di: Kobeissi, Amine, et al.
Pubblicazione: (2026)
Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design
di: Cai, Ruisi, et al.
Pubblicazione: (2024)
di: Cai, Ruisi, et al.
Pubblicazione: (2024)
On the importance of Data Scale in Pretraining Arabic Language Models
di: Ghaddar, Abbas, et al.
Pubblicazione: (2024)
di: Ghaddar, Abbas, et al.
Pubblicazione: (2024)
MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router
di: Xie, Yanyue, et al.
Pubblicazione: (2024)
di: Xie, Yanyue, et al.
Pubblicazione: (2024)
Increasing faithfulness in human-human dialog summarization with Spoken Language Understanding tasks
di: Akani, Eunice, et al.
Pubblicazione: (2024)
di: Akani, Eunice, et al.
Pubblicazione: (2024)
WikiFactDiff: A Large, Realistic, and Temporally Adaptable Dataset for Atomic Factual Knowledge Update in Causal Language Models
di: Khodja, Hichem Ammar, et al.
Pubblicazione: (2024)
di: Khodja, Hichem Ammar, et al.
Pubblicazione: (2024)
Robust Audiovisual Speech Recognition Models with Mixture-of-Experts
di: Wu, Yihan, et al.
Pubblicazione: (2024)
di: Wu, Yihan, et al.
Pubblicazione: (2024)
Glider: Global and Local Instruction-Driven Expert Router
di: Li, Pingzhi, et al.
Pubblicazione: (2024)
di: Li, Pingzhi, et al.
Pubblicazione: (2024)
MMoE: Enhancing Multimodal Models with Mixtures of Multimodal Interaction Experts
di: Yu, Haofei, et al.
Pubblicazione: (2023)
di: Yu, Haofei, et al.
Pubblicazione: (2023)
Factual Knowledge in Language Models: Robustness and Anomalies under Simple Temporal Context Variations
di: Khodja, Hichem Ammar, et al.
Pubblicazione: (2025)
di: Khodja, Hichem Ammar, et al.
Pubblicazione: (2025)
Performance Characterization of Expert Router for Scalable LLM Inference
di: Pichlmeier, Josef, et al.
Pubblicazione: (2024)
di: Pichlmeier, Josef, et al.
Pubblicazione: (2024)
Mixture-of-Experts with Intermediate CTC Supervision for Accented Speech Recognition
di: Lee, Wonjun, et al.
Pubblicazione: (2026)
di: Lee, Wonjun, et al.
Pubblicazione: (2026)
Unveiling Super Experts in Mixture-of-Experts Large Language Models
di: Su, Zunhai, et al.
Pubblicazione: (2025)
di: Su, Zunhai, et al.
Pubblicazione: (2025)
CareMedEval dataset: Evaluating Critical Appraisal and Reasoning in the Biomedical Field
di: Bonzi, Doria, et al.
Pubblicazione: (2025)
di: Bonzi, Doria, et al.
Pubblicazione: (2025)
EUROPA: A Legal Multilingual Keyphrase Generation Dataset
di: Salaün, Olivier, et al.
Pubblicazione: (2024)
di: Salaün, Olivier, et al.
Pubblicazione: (2024)
Mixture of LoRA Experts for Low-Resourced Multi-Accent Automatic Speech Recognition
di: Bagat, Raphaël, et al.
Pubblicazione: (2025)
di: Bagat, Raphaël, et al.
Pubblicazione: (2025)
Cluster-Driven Expert Pruning for Mixture-of-Experts Large Language Models
di: Guo, Hongcheng, et al.
Pubblicazione: (2025)
di: Guo, Hongcheng, et al.
Pubblicazione: (2025)
Mixture of Neuron Experts
di: Cheng, Runxi, et al.
Pubblicazione: (2025)
di: Cheng, Runxi, et al.
Pubblicazione: (2025)
ReGLA: Refining Gated Linear Attention
di: Lu, Peng, et al.
Pubblicazione: (2025)
di: Lu, Peng, et al.
Pubblicazione: (2025)
MoDEM: Mixture of Domain Expert Models
di: Simonds, Toby, et al.
Pubblicazione: (2024)
di: Simonds, Toby, et al.
Pubblicazione: (2024)
CP-Router: An Uncertainty-Aware Router Between LLM and LRM
di: Su, Jiayuan, et al.
Pubblicazione: (2025)
di: Su, Jiayuan, et al.
Pubblicazione: (2025)
Style Mixture of Experts for Expressive Text-To-Speech Synthesis
di: Jawaid, Ahad, et al.
Pubblicazione: (2024)
di: Jawaid, Ahad, et al.
Pubblicazione: (2024)
Enhancing Code-Switching Speech Recognition with LID-Based Collaborative Mixture of Experts Model
di: Huang, Hukai, et al.
Pubblicazione: (2024)
di: Huang, Hukai, et al.
Pubblicazione: (2024)
Chain-of-Experts: Unlocking the Communication Power of Mixture-of-Experts Models
di: Wang, Zihan, et al.
Pubblicazione: (2025)
di: Wang, Zihan, et al.
Pubblicazione: (2025)
EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models
di: Jing, Linglin, et al.
Pubblicazione: (2025)
di: Jing, Linglin, et al.
Pubblicazione: (2025)
Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models
di: Belenki, Lior, et al.
Pubblicazione: (2025)
di: Belenki, Lior, et al.
Pubblicazione: (2025)
MoBiLE: Efficient Mixture-of-Experts Inference on Consumer GPU with Mixture of Big Little Experts
di: Zhao, Yushu, et al.
Pubblicazione: (2025)
di: Zhao, Yushu, et al.
Pubblicazione: (2025)
Mixture of Lookup Experts
di: Jie, Shibo, et al.
Pubblicazione: (2025)
di: Jie, Shibo, et al.
Pubblicazione: (2025)
LABO: Towards Learning Optimal Label Regularization via Bi-level Optimization
di: Lu, Peng, et al.
Pubblicazione: (2023)
di: Lu, Peng, et al.
Pubblicazione: (2023)
DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
di: Dai, Damai, et al.
Pubblicazione: (2024)
di: Dai, Damai, et al.
Pubblicazione: (2024)
Documenti analoghi
-
A linguistically-motivated evaluation methodology for unraveling model's abilities in reading comprehension tasks
di: Antoine, Elie, et al.
Pubblicazione: (2025) -
Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling
di: Ran, Junfeng, et al.
Pubblicazione: (2025) -
Layerwise Recurrent Router for Mixture-of-Experts
di: Qiu, Zihan, et al.
Pubblicazione: (2024) -
On Evaluation Protocols for Data Augmentation in a Limited Data Scenario
di: Piedboeuf, Frédéric, et al.
Pubblicazione: (2024) -
Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss
di: Lv, Ang, et al.
Pubblicazione: (2025)