BlackMamba: Mixture of Experts for State-Space Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Anthony, Quentin, Tokpanov, Yury, Glorioso, Paolo, Millidge, Beren |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference
von: Yao, Jinghan, et al.
Veröffentlicht: (2024)
von: Yao, Jinghan, et al.
Veröffentlicht: (2024)
Zyda-2: a 5 Trillion Token High-Quality Dataset
von: Tokpanov, Yury, et al.
Veröffentlicht: (2024)
von: Tokpanov, Yury, et al.
Veröffentlicht: (2024)
Training Foundation Models on a Full-Stack AMD Platform: Compute, Networking, and System Design
von: Anthony, Quentin, et al.
Veröffentlicht: (2025)
von: Anthony, Quentin, et al.
Veröffentlicht: (2025)
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
FourierMoE: Fourier Mixture-of-Experts Adaptation of Large Language Models
von: Jiang, Juyong, et al.
Veröffentlicht: (2026)
von: Jiang, Juyong, et al.
Veröffentlicht: (2026)
OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models
von: Xue, Fuzhao, et al.
Veröffentlicht: (2024)
von: Xue, Fuzhao, et al.
Veröffentlicht: (2024)
Zamba: A Compact 7B SSM Hybrid Model
von: Glorioso, Paolo, et al.
Veröffentlicht: (2024)
von: Glorioso, Paolo, et al.
Veröffentlicht: (2024)
A Survey on Inference Optimization Techniques for Mixture of Experts Models
von: Liu, Jiacheng, et al.
Veröffentlicht: (2024)
von: Liu, Jiacheng, et al.
Veröffentlicht: (2024)
Optimal Transport Aggregation for Distributed Mixture-of-Experts
von: Chamroukhi, Faïcel, et al.
Veröffentlicht: (2023)
von: Chamroukhi, Faïcel, et al.
Veröffentlicht: (2023)
Utility-Driven Speculative Decoding for Mixture-of-Experts
von: Saxena, Anish, et al.
Veröffentlicht: (2025)
von: Saxena, Anish, et al.
Veröffentlicht: (2025)
Scalable Pretraining of Large Mixture of Experts Language Models on Aurora Super Computer
von: Vooturi, Dharma Teja, et al.
Veröffentlicht: (2026)
von: Vooturi, Dharma Teja, et al.
Veröffentlicht: (2026)
Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts
von: Zhang, Shulai, et al.
Veröffentlicht: (2025)
von: Zhang, Shulai, et al.
Veröffentlicht: (2025)
RaMP: Runtime-Aware Megakernel Polymorphism for Mixture-of-Experts
von: Sharma, Vyom, et al.
Veröffentlicht: (2026)
von: Sharma, Vyom, et al.
Veröffentlicht: (2026)
Scalable Training of Mixture-of-Experts Models with Megatron Core
von: Yan, Zijie, et al.
Veröffentlicht: (2026)
von: Yan, Zijie, et al.
Veröffentlicht: (2026)
Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
Mixture of Experts with Mixture of Precisions for Tuning Quality of Service
von: Imani, HamidReza, et al.
Veröffentlicht: (2024)
von: Imani, HamidReza, et al.
Veröffentlicht: (2024)
From Score Distributions to Balance: Plug-and-Play Mixture-of-Experts Routing
von: Shahout, Rana, et al.
Veröffentlicht: (2025)
von: Shahout, Rana, et al.
Veröffentlicht: (2025)
Mosaic: Data-Free Knowledge Distillation via Mixture-of-Experts for Heterogeneous Distributed Environments
von: Liu, Junming, et al.
Veröffentlicht: (2025)
von: Liu, Junming, et al.
Veröffentlicht: (2025)
The Zamba2 Suite: Technical Report
von: Glorioso, Paolo, et al.
Veröffentlicht: (2024)
von: Glorioso, Paolo, et al.
Veröffentlicht: (2024)
Advancing MoE Efficiency: A Collaboration-Constrained Routing (C2R) Strategy for Better Expert Parallelism Design
von: Zhang, Mohan, et al.
Veröffentlicht: (2025)
von: Zhang, Mohan, et al.
Veröffentlicht: (2025)
Make Every Draft Count: Hidden State based Speculative Decoding
von: Chen, Yuetao, et al.
Veröffentlicht: (2026)
von: Chen, Yuetao, et al.
Veröffentlicht: (2026)
Accelerating MoE Model Inference with Expert Sharding
von: Balmau, Oana, et al.
Veröffentlicht: (2025)
von: Balmau, Oana, et al.
Veröffentlicht: (2025)
Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2024)
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2024)
WDMoE: Wireless Distributed Mixture of Experts for Large Language Models
von: Xue, Nan, et al.
Veröffentlicht: (2024)
von: Xue, Nan, et al.
Veröffentlicht: (2024)
Epistemic Observability in Language Models
von: Mason, Tony, et al.
Veröffentlicht: (2026)
von: Mason, Tony, et al.
Veröffentlicht: (2026)
X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms
von: Yuan, Yueming, et al.
Veröffentlicht: (2025)
von: Yuan, Yueming, et al.
Veröffentlicht: (2025)
Systems and Algorithms for Convolutional Multi-Hybrid Language Models at Scale
von: Ku, Jerome, et al.
Veröffentlicht: (2025)
von: Ku, Jerome, et al.
Veröffentlicht: (2025)
Efficient Deployment of Large Language Models on Resource-constrained Devices
von: Yao, Zhiwei, et al.
Veröffentlicht: (2025)
von: Yao, Zhiwei, et al.
Veröffentlicht: (2025)
PHOTON: Hierarchical Autoregressive Modeling for Lightspeed and Memory-Efficient Language Generation
von: Ichikawa, Yuma, et al.
Veröffentlicht: (2025)
von: Ichikawa, Yuma, et al.
Veröffentlicht: (2025)
Minions: Cost-efficient Collaboration Between On-device and Cloud Language Models
von: Narayan, Avanika, et al.
Veröffentlicht: (2025)
von: Narayan, Avanika, et al.
Veröffentlicht: (2025)
Characterizing and Understanding Energy Footprint and Efficiency of Small Language Model on Edges
von: Islam, Md Romyull, et al.
Veröffentlicht: (2025)
von: Islam, Md Romyull, et al.
Veröffentlicht: (2025)
ReaL: Efficient RLHF Training of Large Language Models with Parameter Reallocation
von: Mei, Zhiyu, et al.
Veröffentlicht: (2024)
von: Mei, Zhiyu, et al.
Veröffentlicht: (2024)
Fisher Information-based Efficient Curriculum Federated Learning with Large Language Models
von: Liu, Ji, et al.
Veröffentlicht: (2024)
von: Liu, Ji, et al.
Veröffentlicht: (2024)
DLoRA: Distributed Parameter-Efficient Fine-Tuning Solution for Large Language Model
von: Gao, Chao, et al.
Veröffentlicht: (2024)
von: Gao, Chao, et al.
Veröffentlicht: (2024)
A Sparsity Predicting Approach for Large Language Models via Activation Pattern Clustering
von: Dhar, Nobel, et al.
Veröffentlicht: (2025)
von: Dhar, Nobel, et al.
Veröffentlicht: (2025)
Distributed Speculative Inference (DSI): Speculation Parallelism for Provably Faster Lossless Language Model Inference
von: Timor, Nadav, et al.
Veröffentlicht: (2024)
von: Timor, Nadav, et al.
Veröffentlicht: (2024)
Toward Sustainable GenAI using Generation Directives for Carbon-Friendly Large Language Model Inference
von: Li, Baolin, et al.
Veröffentlicht: (2024)
von: Li, Baolin, et al.
Veröffentlicht: (2024)
FinGPT-HPC: Efficient Pretraining and Finetuning Large Language Models for Financial Applications with High-Performance Computing
von: Liu, Xiao-Yang, et al.
Veröffentlicht: (2024)
von: Liu, Xiao-Yang, et al.
Veröffentlicht: (2024)
SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State Decoupling
von: Skiadopoulos, Athinagoras, et al.
Veröffentlicht: (2025)
von: Skiadopoulos, Athinagoras, et al.
Veröffentlicht: (2025)
HybridEP: Scaling Expert Parallelism to Cross-Datacenter Scenario via Hybrid Expert/Data Transmission
von: Yang, Weihao, et al.
Veröffentlicht: (2025)
von: Yang, Weihao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference
von: Yao, Jinghan, et al.
Veröffentlicht: (2024) -
Zyda-2: a 5 Trillion Token High-Quality Dataset
von: Tokpanov, Yury, et al.
Veröffentlicht: (2024) -
Training Foundation Models on a Full-Stack AMD Platform: Compute, Networking, and System Design
von: Anthony, Quentin, et al.
Veröffentlicht: (2025) -
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
von: Sun, Weigao, et al.
Veröffentlicht: (2025) -
FourierMoE: Fourier Mixture-of-Experts Adaptation of Large Language Models
von: Jiang, Juyong, et al.
Veröffentlicht: (2026)