FlexMoRE: A Flexible Mixture of Rank-heterogeneous Experts for Efficient Federatedly-trained Large Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Pirchert, Annemette Brok, Nielsen, Jacob, From, Mogens Henrik, Poech, Lukas Galke, Schneider-Kamp, Peter |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DeToNATION: Decoupled Torch Network-Aware Training on Interlinked Online Nodes
por: From, Mogens Henrik, et al.
Publicado: (2025)
por: From, Mogens Henrik, et al.
Publicado: (2025)
The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment
por: Brach, William, et al.
Publicado: (2026)
por: Brach, William, et al.
Publicado: (2026)
Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion
por: Beltoft, Stine Lyngsø, et al.
Publicado: (2026)
por: Beltoft, Stine Lyngsø, et al.
Publicado: (2026)
SDUs DAISY: A Benchmark for Danish Culture
por: Nielsen, Jacob, et al.
Publicado: (2026)
por: Nielsen, Jacob, et al.
Publicado: (2026)
Continual Quantization-Aware Pre-Training: When to transition from 16-bit to 1.58-bit pre-training for BitNet language models?
por: Nielsen, Jacob, et al.
Publicado: (2025)
por: Nielsen, Jacob, et al.
Publicado: (2025)
Training Language Models to Use Prolog as a Tool
por: Mellgren, Niklas, et al.
Publicado: (2025)
por: Mellgren, Niklas, et al.
Publicado: (2025)
Confidence and Calibration of Activation Oracles for Reliable Interpretation of Language Model Internals
por: Torrielli, Federico, et al.
Publicado: (2026)
por: Torrielli, Federico, et al.
Publicado: (2026)
When are 1.58 bits enough? A Bottom-up Exploration of BitNet Quantization
por: Nielsen, Jacob, et al.
Publicado: (2024)
por: Nielsen, Jacob, et al.
Publicado: (2024)
Isolating Culture Neurons in Multilingual Large Language Models
por: Namazifard, Danial, et al.
Publicado: (2025)
por: Namazifard, Danial, et al.
Publicado: (2025)
DaLA: Danish Linguistic Acceptability Evaluation Guided by Real World Errors
por: Barmina, Gianluca, et al.
Publicado: (2025)
por: Barmina, Gianluca, et al.
Publicado: (2025)
Chain of Summaries: Summarization Through Iterative Questioning
por: Brach, William, et al.
Publicado: (2025)
por: Brach, William, et al.
Publicado: (2025)
Flex-MoE: Modeling Arbitrary Modality Combination via the Flexible Mixture-of-Experts
por: Yun, Sukwon, et al.
Publicado: (2024)
por: Yun, Sukwon, et al.
Publicado: (2024)
MoRE: A Mixture of Low-Rank Experts for Adaptive Multi-Task Learning
por: Zhang, Dacao, et al.
Publicado: (2025)
por: Zhang, Dacao, et al.
Publicado: (2025)
BitNet b1.58 Reloaded: State-of-the-art Performance Also on Smaller Networks
por: Nielsen, Jacob, et al.
Publicado: (2024)
por: Nielsen, Jacob, et al.
Publicado: (2024)
SommBench: Assessing Sommelier Expertise of Language Models
por: Brach, William, et al.
Publicado: (2026)
por: Brach, William, et al.
Publicado: (2026)
ChronoMedKG: A Temporally-Grounded Biomedical Knowledge Graph and Benchmark for Clinical Reasoning
por: Ahmed, Md Shamim, et al.
Publicado: (2026)
por: Ahmed, Md Shamim, et al.
Publicado: (2026)
S'MoRE: Structural Mixture of Residual Experts for Parameter-Efficient LLM Fine-tuning
por: Zeng, Hanqing, et al.
Publicado: (2025)
por: Zeng, Hanqing, et al.
Publicado: (2025)
MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-Experts
por: Gao, Jingnan, et al.
Publicado: (2025)
por: Gao, Jingnan, et al.
Publicado: (2025)
GraphMoRE: Mitigating Topological Heterogeneity via Mixture of Riemannian Experts
por: Guo, Zihao, et al.
Publicado: (2024)
por: Guo, Zihao, et al.
Publicado: (2024)
MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts
por: Lin, Xi Victoria, et al.
Publicado: (2024)
por: Lin, Xi Victoria, et al.
Publicado: (2024)
MoRE-LLM: Mixture of Rule Experts Guided by a Large Language Model
por: Koebler, Alexander, et al.
Publicado: (2025)
por: Koebler, Alexander, et al.
Publicado: (2025)
MoRE: Mixture of Residual Experts for Humanoid Lifelike Gaits Learning on Complex Terrains
por: Wang, Dewei, et al.
Publicado: (2025)
por: Wang, Dewei, et al.
Publicado: (2025)
Mixture of Robust Experts (MoRE):A Robust Denoising Method towards multiple perturbations
por: Cheng, Hao, et al.
Publicado: (2021)
por: Cheng, Hao, et al.
Publicado: (2021)
FlexEControl: Flexible and Efficient Multimodal Control for Text-to-Image Generation
por: He, Xuehai, et al.
Publicado: (2024)
por: He, Xuehai, et al.
Publicado: (2024)
ALoRE: Efficient Visual Adaptation via Aggregating Low Rank Experts
por: Du, Sinan, et al.
Publicado: (2024)
por: Du, Sinan, et al.
Publicado: (2024)
FlexLoRA: Entropy-Guided Flexible Low-Rank Adaptation
por: Liu, Muqing, et al.
Publicado: (2026)
por: Liu, Muqing, et al.
Publicado: (2026)
MobileMoE: Scaling On-Device Mixture of Experts
por: Chen, Yanbei, et al.
Publicado: (2026)
por: Chen, Yanbei, et al.
Publicado: (2026)
MoBiLE: Efficient Mixture-of-Experts Inference on Consumer GPU with Mixture of Big Little Experts
por: Zhao, Yushu, et al.
Publicado: (2025)
por: Zhao, Yushu, et al.
Publicado: (2025)
chembl/EnsembleFlex: Published version of EnsembleFlex
por: Melanie Schneider
Publicado: (2025)
por: Melanie Schneider
Publicado: (2025)
The Provenance Gap in Clinical AI: Evidence-Traceable Temporal Knowledge Graphs for Rare Disease Reasoning
por: Ahmed, Md Shamim, et al.
Publicado: (2026)
por: Ahmed, Md Shamim, et al.
Publicado: (2026)
Super-additive Cooperation in Language Model Agents
por: Tonini, Filippo, et al.
Publicado: (2025)
por: Tonini, Filippo, et al.
Publicado: (2025)
Not Everything That Counts Can Be Counted: A Case for Safe Qualitative AI
por: Beltoft, Stine, et al.
Publicado: (2025)
por: Beltoft, Stine, et al.
Publicado: (2025)
Learning and communication pressures in neural networks: Lessons from emergent communication
por: Galke, Lukas, et al.
Publicado: (2024)
por: Galke, Lukas, et al.
Publicado: (2024)
Decidability Issues for Petri Nets -- a survey
por: Esparza, Javier, et al.
Publicado: (2024)
por: Esparza, Javier, et al.
Publicado: (2024)
FLEX-MoE: Federated Mixture-of-Experts with Load-balanced Expert Assignment for Edge Computing
por: Zhang, Boyang, et al.
Publicado: (2025)
por: Zhang, Boyang, et al.
Publicado: (2025)
MoRE: A Mixture-of-Experts-Based Task-Adaptive End-to-End Network for Multimodal MRI Reconstruction
por: Li, Yuyang, et al.
Publicado: (2026)
por: Li, Yuyang, et al.
Publicado: (2026)
MoRE-Brain: Routed Mixture of Experts for Interpretable and Generalizable Cross-Subject fMRI Visual Decoding
por: Wei, Yuxiang, et al.
Publicado: (2025)
por: Wei, Yuxiang, et al.
Publicado: (2025)
PreMoE: Proactive Inference for Efficient Mixture-of-Experts
por: Pei, Zehua, et al.
Publicado: (2025)
por: Pei, Zehua, et al.
Publicado: (2025)
OrdMoE: Preference Alignment via Hierarchical Expert Group Ranking in Multimodal Mixture-of-Experts LLMs
por: Gao, Yuting, et al.
Publicado: (2025)
por: Gao, Yuting, et al.
Publicado: (2025)
Elastic Mixture of Rank-Wise Experts for Knowledge Reuse in Federated Fine-Tuning
por: Wu, Yebo, et al.
Publicado: (2025)
por: Wu, Yebo, et al.
Publicado: (2025)
Ejemplares similares
-
DeToNATION: Decoupled Torch Network-Aware Training on Interlinked Online Nodes
por: From, Mogens Henrik, et al.
Publicado: (2025) -
The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment
por: Brach, William, et al.
Publicado: (2026) -
Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion
por: Beltoft, Stine Lyngsø, et al.
Publicado: (2026) -
SDUs DAISY: A Benchmark for Danish Culture
por: Nielsen, Jacob, et al.
Publicado: (2026) -
Continual Quantization-Aware Pre-Training: When to transition from 16-bit to 1.58-bit pre-training for BitNet language models?
por: Nielsen, Jacob, et al.
Publicado: (2025)