Path-Constrained Mixture-of-Experts
Fuente:
arXiv
Saved in:
| Main Authors: | Gu, Zijin, Likhomanenko, Tatiana, Thilak, Vimal, Ramapuram, Jason, Jaitly, Navdeep |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Omni-Router: Sharing Routing Decisions in Sparse Mixture-of-Experts for Speech Recognition
by: Gu, Zijin, et al.
Published: (2025)
by: Gu, Zijin, et al.
Published: (2025)
SpeakStream: Streaming Text-to-Speech with Interleaved Data
by: Bai, Richard He, et al.
Published: (2025)
by: Bai, Richard He, et al.
Published: (2025)
Scaling Properties of Continuous Diffusion Spoken Language Models
by: Ramapuram, Jason, et al.
Published: (2026)
by: Ramapuram, Jason, et al.
Published: (2026)
Revisiting ASR Error Correction with Specialized Models
by: Gu, Zijin, et al.
Published: (2024)
by: Gu, Zijin, et al.
Published: (2024)
Theory, Analysis, and Best Practices for Sigmoid Self-Attention
by: Ramapuram, Jason, et al.
Published: (2024)
by: Ramapuram, Jason, et al.
Published: (2024)
ChipChat: Low-Latency Cascaded Conversational Agent in MLX
by: Likhomanenko, Tatiana, et al.
Published: (2025)
by: Likhomanenko, Tatiana, et al.
Published: (2025)
Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models
by: Abnar, Samira, et al.
Published: (2025)
by: Abnar, Samira, et al.
Published: (2025)
dMel: Speech Tokenization made Simple
by: Bai, Richard He, et al.
Published: (2024)
by: Bai, Richard He, et al.
Published: (2024)
What Makes the Preferred Thinking Direction for LLMs in Multiple-choice Questions?
by: Zhang, Yizhe, et al.
Published: (2025)
by: Zhang, Yizhe, et al.
Published: (2025)
Enhancing JEPAs with Spatial Conditioning: Robust and Efficient Representation Learning
by: Littwin, Etai, et al.
Published: (2024)
by: Littwin, Etai, et al.
Published: (2024)
Probing the Multi-turn Planning Capabilities of LLMs via 20 Question Games
by: Zhang, Yizhe, et al.
Published: (2023)
by: Zhang, Yizhe, et al.
Published: (2023)
Matryoshka Diffusion Models
by: Gu, Jiatao, et al.
Published: (2023)
by: Gu, Jiatao, et al.
Published: (2023)
Towards Automatic Assessment of Self-Supervised Speech Models using Rank
by: Aldeneh, Zakaria, et al.
Published: (2024)
by: Aldeneh, Zakaria, et al.
Published: (2024)
Closing the Gap Between Text and Speech Understanding in LLMs
by: Cuervo, Santiago, et al.
Published: (2025)
by: Cuervo, Santiago, et al.
Published: (2025)
Text-Conditional JEPA for Learning Semantically Rich Visual Representations
by: Huang, Chen, et al.
Published: (2026)
by: Huang, Chen, et al.
Published: (2026)
Continuously Augmented Discrete Diffusion model for Categorical Generative Modeling
by: Zheng, Huangjie, et al.
Published: (2025)
by: Zheng, Huangjie, et al.
Published: (2025)
Unsupervised ASR via Cross-Lingual Pseudo-Labeling
by: Likhomanenko, Tatiana, et al.
Published: (2023)
by: Likhomanenko, Tatiana, et al.
Published: (2023)
SimpleFold: Folding Proteins is Simpler than You Think
by: Wang, Yuyang, et al.
Published: (2025)
by: Wang, Yuyang, et al.
Published: (2025)
Divide-or-Conquer? Which Part Should You Distill Your LLM?
by: Wu, Zhuofeng, et al.
Published: (2024)
by: Wu, Zhuofeng, et al.
Published: (2024)
How JEPA Avoids Noisy Features: The Implicit Bias of Deep Linear Self Distillation Networks
by: Littwin, Etai, et al.
Published: (2024)
by: Littwin, Etai, et al.
Published: (2024)
Prediction-powered Inference by Mixture of Experts
by: Gu, Yanwu, et al.
Published: (2026)
by: Gu, Yanwu, et al.
Published: (2026)
Target Concrete Score Matching: A Holistic Framework for Discrete Diffusion
by: Zhang, Ruixiang, et al.
Published: (2025)
by: Zhang, Ruixiang, et al.
Published: (2025)
Improving GFlowNets for Text-to-Image Diffusion Alignment
by: Zhang, Dinghuai, et al.
Published: (2024)
by: Zhang, Dinghuai, et al.
Published: (2024)
Swallowing the Bitter Pill: Simplified Scalable Conformer Generation
by: Wang, Yuyang, et al.
Published: (2023)
by: Wang, Yuyang, et al.
Published: (2023)
Aggregate-and-Adapt Natural Language Prompts for Downstream Generalization of CLIP
by: Huang, Chen, et al.
Published: (2024)
by: Huang, Chen, et al.
Published: (2024)
KGLens: Towards Efficient and Effective Knowledge Probing of Large Language Models with Knowledge Graphs
by: Zheng, Shangshang, et al.
Published: (2023)
by: Zheng, Shangshang, et al.
Published: (2023)
BuddyMoE: Exploiting Expert Redundancy to Accelerate Memory-Constrained Mixture-of-Experts Inference
by: Wang, Yun, et al.
Published: (2025)
by: Wang, Yun, et al.
Published: (2025)
DART: Denoising Autoregressive Transformer for Scalable Text-to-Image Generation
by: Gu, Jiatao, et al.
Published: (2024)
by: Gu, Jiatao, et al.
Published: (2024)
PolyBlocks: A Compiler Infrastructure for AI Chips and Programming Frameworks
by: Bondhugula, Uday, et al.
Published: (2026)
by: Bondhugula, Uday, et al.
Published: (2026)
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
by: Li, Lujun, et al.
Published: (2025)
by: Li, Lujun, et al.
Published: (2025)
Flexible Language Modeling in Continuous Space with Transformer-based Autoregressive Flows
by: Zhang, Ruixiang, et al.
Published: (2025)
by: Zhang, Ruixiang, et al.
Published: (2025)
Revisiting the Scaling Properties of Downstream Metrics in Large Language Model Training
by: Krajewski, Jakub, et al.
Published: (2025)
by: Krajewski, Jakub, et al.
Published: (2025)
Learning More Generalized Experts by Merging Experts in Mixture-of-Experts
by: Park, Sejik
Published: (2024)
by: Park, Sejik
Published: (2024)
MoEEdit: Efficient and Routing-Stable Knowledge Editing for Mixture-of-Experts LLMs
by: Gu, Yupu, et al.
Published: (2026)
by: Gu, Yupu, et al.
Published: (2026)
Mechanisms of Multimodal Synchronization: Insights from Decoder-Based Video-Text-to-Speech Synthesis
by: Gupta, Akshita, et al.
Published: (2024)
by: Gupta, Akshita, et al.
Published: (2024)
MoDE: A Mixture-of-Experts Model with Mutual Distillation among the Experts
by: Xie, Zhitian, et al.
Published: (2024)
by: Xie, Zhitian, et al.
Published: (2024)
Nearly Minimax Optimal Regret for Learning Linear Mixture Stochastic Shortest Path
by: Di, Qiwei, et al.
Published: (2024)
by: Di, Qiwei, et al.
Published: (2024)
Expert Merging in Sparse Mixture of Experts with Nash Bargaining
by: Nguyen, Dung V., et al.
Published: (2025)
by: Nguyen, Dung V., et al.
Published: (2025)
$μ$-Parametrization for Mixture of Experts
by: Małaśnicki, Jan, et al.
Published: (2025)
by: Małaśnicki, Jan, et al.
Published: (2025)
TileQ: Efficient Low-Rank Quantization of Mixture-of-Experts with 2D Tiling
by: Gu, Hongyaoxing, et al.
Published: (2026)
by: Gu, Hongyaoxing, et al.
Published: (2026)
Similar Items
-
Omni-Router: Sharing Routing Decisions in Sparse Mixture-of-Experts for Speech Recognition
by: Gu, Zijin, et al.
Published: (2025) -
SpeakStream: Streaming Text-to-Speech with Interleaved Data
by: Bai, Richard He, et al.
Published: (2025) -
Scaling Properties of Continuous Diffusion Spoken Language Models
by: Ramapuram, Jason, et al.
Published: (2026) -
Revisiting ASR Error Correction with Specialized Models
by: Gu, Zijin, et al.
Published: (2024) -
Theory, Analysis, and Best Practices for Sigmoid Self-Attention
by: Ramapuram, Jason, et al.
Published: (2024)