MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Xi Victoria, Shrivastava, Akshat, Luo, Liang, Iyer, Srinivasan, Lewis, Mike, Ghosh, Gargi, Zettlemoyer, Luke, Aghajanyan, Armen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving MoE Compute Efficiency by Composing Weight and Data Sparsity
by: Kilian, Maciej, et al.
Published: (2026)
by: Kilian, Maciej, et al.
Published: (2026)
Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models
by: Liang, Weixin, et al.
Published: (2024)
by: Liang, Weixin, et al.
Published: (2024)
SkillRater: Untangling Capabilities in Multimodal Data
by: Sahi, Naveen, et al.
Published: (2026)
by: Sahi, Naveen, et al.
Published: (2026)
When Worse is Better: Navigating the compression-generation tradeoff in visual tokenization
by: Ramanujan, Vivek, et al.
Published: (2024)
by: Ramanujan, Vivek, et al.
Published: (2024)
Lory: Fully Differentiable Mixture-of-Experts for Autoregressive Language Model Pre-training
by: Zhong, Zexuan, et al.
Published: (2024)
by: Zhong, Zexuan, et al.
Published: (2024)
Compute Optimal Tokenization
by: Limisiewicz, Tomasz, et al.
Published: (2026)
by: Limisiewicz, Tomasz, et al.
Published: (2026)
MoMa-Pos: An Efficient Object-Kinematic-Aware Base Placement Optimization Framework for Mobile Manipulation
by: Shao, Beichen, et al.
Published: (2024)
by: Shao, Beichen, et al.
Published: (2024)
Fast Byte Latent Transformer
by: Kallini, Julie, et al.
Published: (2026)
by: Kallini, Julie, et al.
Published: (2026)
MoMa: A Modular Deep Learning Framework for Material Property Prediction
by: Wang, Botian, et al.
Published: (2025)
by: Wang, Botian, et al.
Published: (2025)
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition
by: Yang, Yuhuan, et al.
Published: (2025)
by: Yang, Yuhuan, et al.
Published: (2025)
Text Quality-Based Pruning for Efficient Training of Language Models
by: Sharma, Vasu, et al.
Published: (2024)
by: Sharma, Vasu, et al.
Published: (2024)
Latent Speech-Text Transformer
by: Lu, Yen-Ju, et al.
Published: (2025)
by: Lu, Yen-Ju, et al.
Published: (2025)
Sketch-MoMa: Teleoperation for Mobile Manipulator via Interpretation of Hand-Drawn Sketches
by: Tanada, Kosei, et al.
Published: (2024)
by: Tanada, Kosei, et al.
Published: (2024)
Slicing and Dicing: Configuring Optimal Mixtures of Experts
by: Li, Margaret, et al.
Published: (2026)
by: Li, Margaret, et al.
Published: (2026)
Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity
by: Liang, Weixin, et al.
Published: (2025)
by: Liang, Weixin, et al.
Published: (2025)
AIRoA MoMa Dataset: A Large-Scale Hierarchical Dataset for Mobile Manipulation
by: Takanami, Ryosuke, et al.
Published: (2025)
by: Takanami, Ryosuke, et al.
Published: (2025)
CoSMoEs: Compact Sparse Mixture of Experts
by: Huber, Patrick, et al.
Published: (2025)
by: Huber, Patrick, et al.
Published: (2025)
MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile Manipulation
by: Zhang, Pingrui, et al.
Published: (2025)
by: Zhang, Pingrui, et al.
Published: (2025)
MoST: Mixing Speech and Text with Modality-Aware Mixture of Experts
by: Lou, Yuxuan, et al.
Published: (2026)
by: Lou, Yuxuan, et al.
Published: (2026)
Byte Latent Transformer: Patches Scale Better Than Tokens
by: Pagnoni, Artidoro, et al.
Published: (2024)
by: Pagnoni, Artidoro, et al.
Published: (2024)
Symphony-MoE: Harmonizing Disparate Pre-trained Models into a Coherent Mixture-of-Experts
by: Wang, Qi, et al.
Published: (2025)
by: Wang, Qi, et al.
Published: (2025)
LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training
by: Zhu, Tong, et al.
Published: (2024)
by: Zhu, Tong, et al.
Published: (2024)
eMoE: Task-aware Memory Efficient Mixture-of-Experts-Based (MoE) Model Inference
by: Tairin, Suraiya, et al.
Published: (2025)
by: Tairin, Suraiya, et al.
Published: (2025)
PreMoE: Proactive Inference for Efficient Mixture-of-Experts
by: Pei, Zehua, et al.
Published: (2025)
by: Pei, Zehua, et al.
Published: (2025)
FuseMoE: Mixture-of-Experts Transformers for Fleximodal Fusion
by: Han, Xing, et al.
Published: (2024)
by: Han, Xing, et al.
Published: (2024)
Memory Layers at Scale
by: Berges, Vincent-Pierre, et al.
Published: (2024)
by: Berges, Vincent-Pierre, et al.
Published: (2024)
EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoE
by: Chen, Junyi, et al.
Published: (2023)
by: Chen, Junyi, et al.
Published: (2023)
Facet-Aware Multi-Head Mixture-of-Experts Model with Text-Enhanced Pre-training for Sequential Recommendation
by: Liu, Mingrui, et al.
Published: (2026)
by: Liu, Mingrui, et al.
Published: (2026)
TiMoE: Time-Aware Mixture of Language Experts
by: Faro, Robin, et al.
Published: (2025)
by: Faro, Robin, et al.
Published: (2025)
MoDE: CLIP Data Experts via Clustering
by: Ma, Jiawei, et al.
Published: (2024)
by: Ma, Jiawei, et al.
Published: (2024)
The weighted Bergman spaces and complex reflection groups
by: Ghosh, Gargi
Published: (2021)
by: Ghosh, Gargi
Published: (2021)
Small Molecule Optimization with Large Language Models
by: Guevorguian, Philipp, et al.
Published: (2024)
by: Guevorguian, Philipp, et al.
Published: (2024)
MoPE: Mixture of Prompt Experts for Parameter-Efficient and Scalable Multimodal Fusion
by: Jiang, Ruixiang, et al.
Published: (2024)
by: Jiang, Ruixiang, et al.
Published: (2024)
IMA-MoE: An Interpretable Modality-Aware Mixture-of-Experts Framework for Characterizing the Neurobiological Signatures of Binge Eating Disorder
by: Zhao, Lin, et al.
Published: (2026)
by: Zhao, Lin, et al.
Published: (2026)
Continual Learning via Sparse Memory Finetuning
by: Lin, Jessy, et al.
Published: (2025)
by: Lin, Jessy, et al.
Published: (2025)
Recycling the Web: A Method to Enhance Pre-training Data Quality and Quantity for Language Models
by: Nguyen, Thao, et al.
Published: (2025)
by: Nguyen, Thao, et al.
Published: (2025)
PWC-MoE: Privacy-Aware Wireless Collaborative Mixture of Experts
by: Su, Yang, et al.
Published: (2025)
by: Su, Yang, et al.
Published: (2025)
OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models
by: Xue, Fuzhao, et al.
Published: (2024)
by: Xue, Fuzhao, et al.
Published: (2024)
Hierarchical Time-Aware Mixture of Experts for Multi-Modal Sequential Recommendation
by: Zhang, Shengzhe, et al.
Published: (2025)
by: Zhang, Shengzhe, et al.
Published: (2025)
Flex-MoE: Modeling Arbitrary Modality Combination via the Flexible Mixture-of-Experts
by: Yun, Sukwon, et al.
Published: (2024)
by: Yun, Sukwon, et al.
Published: (2024)
Similar Items
-
Improving MoE Compute Efficiency by Composing Weight and Data Sparsity
by: Kilian, Maciej, et al.
Published: (2026) -
Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models
by: Liang, Weixin, et al.
Published: (2024) -
SkillRater: Untangling Capabilities in Multimodal Data
by: Sahi, Naveen, et al.
Published: (2026) -
When Worse is Better: Navigating the compression-generation tradeoff in visual tokenization
by: Ramanujan, Vivek, et al.
Published: (2024) -
Lory: Fully Differentiable Mixture-of-Experts for Autoregressive Language Model Pre-training
by: Zhong, Zexuan, et al.
Published: (2024)