MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Xi Victoria, Shrivastava, Akshat, Luo, Liang, Iyer, Srinivasan, Lewis, Mike, Ghosh, Gargi, Zettlemoyer, Luke, Aghajanyan, Armen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Improving MoE Compute Efficiency by Composing Weight and Data Sparsity
von: Kilian, Maciej, et al.
Veröffentlicht: (2026)
von: Kilian, Maciej, et al.
Veröffentlicht: (2026)
Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models
von: Liang, Weixin, et al.
Veröffentlicht: (2024)
von: Liang, Weixin, et al.
Veröffentlicht: (2024)
SkillRater: Untangling Capabilities in Multimodal Data
von: Sahi, Naveen, et al.
Veröffentlicht: (2026)
von: Sahi, Naveen, et al.
Veröffentlicht: (2026)
When Worse is Better: Navigating the compression-generation tradeoff in visual tokenization
von: Ramanujan, Vivek, et al.
Veröffentlicht: (2024)
von: Ramanujan, Vivek, et al.
Veröffentlicht: (2024)
Lory: Fully Differentiable Mixture-of-Experts for Autoregressive Language Model Pre-training
von: Zhong, Zexuan, et al.
Veröffentlicht: (2024)
von: Zhong, Zexuan, et al.
Veröffentlicht: (2024)
Compute Optimal Tokenization
von: Limisiewicz, Tomasz, et al.
Veröffentlicht: (2026)
von: Limisiewicz, Tomasz, et al.
Veröffentlicht: (2026)
MoMa-Pos: An Efficient Object-Kinematic-Aware Base Placement Optimization Framework for Mobile Manipulation
von: Shao, Beichen, et al.
Veröffentlicht: (2024)
von: Shao, Beichen, et al.
Veröffentlicht: (2024)
Fast Byte Latent Transformer
von: Kallini, Julie, et al.
Veröffentlicht: (2026)
von: Kallini, Julie, et al.
Veröffentlicht: (2026)
MoMa: A Modular Deep Learning Framework for Material Property Prediction
von: Wang, Botian, et al.
Veröffentlicht: (2025)
von: Wang, Botian, et al.
Veröffentlicht: (2025)
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition
von: Yang, Yuhuan, et al.
Veröffentlicht: (2025)
von: Yang, Yuhuan, et al.
Veröffentlicht: (2025)
Text Quality-Based Pruning for Efficient Training of Language Models
von: Sharma, Vasu, et al.
Veröffentlicht: (2024)
von: Sharma, Vasu, et al.
Veröffentlicht: (2024)
Latent Speech-Text Transformer
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2025)
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2025)
Sketch-MoMa: Teleoperation for Mobile Manipulator via Interpretation of Hand-Drawn Sketches
von: Tanada, Kosei, et al.
Veröffentlicht: (2024)
von: Tanada, Kosei, et al.
Veröffentlicht: (2024)
Slicing and Dicing: Configuring Optimal Mixtures of Experts
von: Li, Margaret, et al.
Veröffentlicht: (2026)
von: Li, Margaret, et al.
Veröffentlicht: (2026)
Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity
von: Liang, Weixin, et al.
Veröffentlicht: (2025)
von: Liang, Weixin, et al.
Veröffentlicht: (2025)
AIRoA MoMa Dataset: A Large-Scale Hierarchical Dataset for Mobile Manipulation
von: Takanami, Ryosuke, et al.
Veröffentlicht: (2025)
von: Takanami, Ryosuke, et al.
Veröffentlicht: (2025)
CoSMoEs: Compact Sparse Mixture of Experts
von: Huber, Patrick, et al.
Veröffentlicht: (2025)
von: Huber, Patrick, et al.
Veröffentlicht: (2025)
MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile Manipulation
von: Zhang, Pingrui, et al.
Veröffentlicht: (2025)
von: Zhang, Pingrui, et al.
Veröffentlicht: (2025)
MoST: Mixing Speech and Text with Modality-Aware Mixture of Experts
von: Lou, Yuxuan, et al.
Veröffentlicht: (2026)
von: Lou, Yuxuan, et al.
Veröffentlicht: (2026)
Byte Latent Transformer: Patches Scale Better Than Tokens
von: Pagnoni, Artidoro, et al.
Veröffentlicht: (2024)
von: Pagnoni, Artidoro, et al.
Veröffentlicht: (2024)
Symphony-MoE: Harmonizing Disparate Pre-trained Models into a Coherent Mixture-of-Experts
von: Wang, Qi, et al.
Veröffentlicht: (2025)
von: Wang, Qi, et al.
Veröffentlicht: (2025)
LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training
von: Zhu, Tong, et al.
Veröffentlicht: (2024)
von: Zhu, Tong, et al.
Veröffentlicht: (2024)
eMoE: Task-aware Memory Efficient Mixture-of-Experts-Based (MoE) Model Inference
von: Tairin, Suraiya, et al.
Veröffentlicht: (2025)
von: Tairin, Suraiya, et al.
Veröffentlicht: (2025)
PreMoE: Proactive Inference for Efficient Mixture-of-Experts
von: Pei, Zehua, et al.
Veröffentlicht: (2025)
von: Pei, Zehua, et al.
Veröffentlicht: (2025)
FuseMoE: Mixture-of-Experts Transformers for Fleximodal Fusion
von: Han, Xing, et al.
Veröffentlicht: (2024)
von: Han, Xing, et al.
Veröffentlicht: (2024)
Memory Layers at Scale
von: Berges, Vincent-Pierre, et al.
Veröffentlicht: (2024)
von: Berges, Vincent-Pierre, et al.
Veröffentlicht: (2024)
EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoE
von: Chen, Junyi, et al.
Veröffentlicht: (2023)
von: Chen, Junyi, et al.
Veröffentlicht: (2023)
Facet-Aware Multi-Head Mixture-of-Experts Model with Text-Enhanced Pre-training for Sequential Recommendation
von: Liu, Mingrui, et al.
Veröffentlicht: (2026)
von: Liu, Mingrui, et al.
Veröffentlicht: (2026)
TiMoE: Time-Aware Mixture of Language Experts
von: Faro, Robin, et al.
Veröffentlicht: (2025)
von: Faro, Robin, et al.
Veröffentlicht: (2025)
MoDE: CLIP Data Experts via Clustering
von: Ma, Jiawei, et al.
Veröffentlicht: (2024)
von: Ma, Jiawei, et al.
Veröffentlicht: (2024)
The weighted Bergman spaces and complex reflection groups
von: Ghosh, Gargi
Veröffentlicht: (2021)
von: Ghosh, Gargi
Veröffentlicht: (2021)
Small Molecule Optimization with Large Language Models
von: Guevorguian, Philipp, et al.
Veröffentlicht: (2024)
von: Guevorguian, Philipp, et al.
Veröffentlicht: (2024)
MoPE: Mixture of Prompt Experts for Parameter-Efficient and Scalable Multimodal Fusion
von: Jiang, Ruixiang, et al.
Veröffentlicht: (2024)
von: Jiang, Ruixiang, et al.
Veröffentlicht: (2024)
IMA-MoE: An Interpretable Modality-Aware Mixture-of-Experts Framework for Characterizing the Neurobiological Signatures of Binge Eating Disorder
von: Zhao, Lin, et al.
Veröffentlicht: (2026)
von: Zhao, Lin, et al.
Veröffentlicht: (2026)
Continual Learning via Sparse Memory Finetuning
von: Lin, Jessy, et al.
Veröffentlicht: (2025)
von: Lin, Jessy, et al.
Veröffentlicht: (2025)
Recycling the Web: A Method to Enhance Pre-training Data Quality and Quantity for Language Models
von: Nguyen, Thao, et al.
Veröffentlicht: (2025)
von: Nguyen, Thao, et al.
Veröffentlicht: (2025)
PWC-MoE: Privacy-Aware Wireless Collaborative Mixture of Experts
von: Su, Yang, et al.
Veröffentlicht: (2025)
von: Su, Yang, et al.
Veröffentlicht: (2025)
OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models
von: Xue, Fuzhao, et al.
Veröffentlicht: (2024)
von: Xue, Fuzhao, et al.
Veröffentlicht: (2024)
Hierarchical Time-Aware Mixture of Experts for Multi-Modal Sequential Recommendation
von: Zhang, Shengzhe, et al.
Veröffentlicht: (2025)
von: Zhang, Shengzhe, et al.
Veröffentlicht: (2025)
Flex-MoE: Modeling Arbitrary Modality Combination via the Flexible Mixture-of-Experts
von: Yun, Sukwon, et al.
Veröffentlicht: (2024)
von: Yun, Sukwon, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Improving MoE Compute Efficiency by Composing Weight and Data Sparsity
von: Kilian, Maciej, et al.
Veröffentlicht: (2026) -
Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models
von: Liang, Weixin, et al.
Veröffentlicht: (2024) -
SkillRater: Untangling Capabilities in Multimodal Data
von: Sahi, Naveen, et al.
Veröffentlicht: (2026) -
When Worse is Better: Navigating the compression-generation tradeoff in visual tokenization
von: Ramanujan, Vivek, et al.
Veröffentlicht: (2024) -
Lory: Fully Differentiable Mixture-of-Experts for Autoregressive Language Model Pre-training
von: Zhong, Zexuan, et al.
Veröffentlicht: (2024)