MoM: Linear Sequence Modeling with Mixture-of-Memories
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Du, Jusen, Sun, Weigao, Lan, Disen, Hu, Jiaxi, Cheng, Yu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Liger: Linearizing Large Language Models to Gated Recurrent Structures
von: Lan, Disen, et al.
Veröffentlicht: (2025)
von: Lan, Disen, et al.
Veröffentlicht: (2025)
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
Native Hybrid Attention for Efficient Sequence Modeling
von: Du, Jusen, et al.
Veröffentlicht: (2025)
von: Du, Jusen, et al.
Veröffentlicht: (2025)
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
Comba: Improving Bilinear RNNs with Closed-loop Control
von: Hu, Jiaxi, et al.
Veröffentlicht: (2025)
von: Hu, Jiaxi, et al.
Veröffentlicht: (2025)
Speed Always Wins: A Survey on Efficient Architectures for Large Language Models
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
Time-SSM: Simplifying and Unifying State Space Models for Time Series Forecasting
von: Hu, Jiaxi, et al.
Veröffentlicht: (2024)
von: Hu, Jiaxi, et al.
Veröffentlicht: (2024)
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
von: Ludziejewski, Jan, et al.
Veröffentlicht: (2025)
von: Ludziejewski, Jan, et al.
Veröffentlicht: (2025)
MoKA: Mixture of Kronecker Adapters
von: Sadeghi, Mohammadreza, et al.
Veröffentlicht: (2025)
von: Sadeghi, Mohammadreza, et al.
Veröffentlicht: (2025)
QuantMoE-Bench: Examining Post-Training Quantization for Mixture-of-Experts
von: Li, Pingzhi, et al.
Veröffentlicht: (2024)
von: Li, Pingzhi, et al.
Veröffentlicht: (2024)
LeMoLE: LLM-Enhanced Mixture of Linear Experts for Time Series Forecasting
von: Zhang, Lingzheng, et al.
Veröffentlicht: (2024)
von: Zhang, Lingzheng, et al.
Veröffentlicht: (2024)
MeMo: Memory as a Model
von: Quek, Ryan Wei Heng, et al.
Veröffentlicht: (2026)
von: Quek, Ryan Wei Heng, et al.
Veröffentlicht: (2026)
Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
von: Gu, Naibin, et al.
Veröffentlicht: (2025)
von: Gu, Naibin, et al.
Veröffentlicht: (2025)
MOM: Memory-Efficient Offloaded Mini-Sequence Inference for Long Context Language Models
von: Zhang, Junyang, et al.
Veröffentlicht: (2025)
von: Zhang, Junyang, et al.
Veröffentlicht: (2025)
Mini-batch Coresets for Memory-efficient Language Model Training on Data Mixtures
von: Nguyen, Dang, et al.
Veröffentlicht: (2024)
von: Nguyen, Dang, et al.
Veröffentlicht: (2024)
MoIN: Mixture of Introvert Experts to Upcycle an LLM
von: Tejankar, Ajinkya, et al.
Veröffentlicht: (2024)
von: Tejankar, Ajinkya, et al.
Veröffentlicht: (2024)
MoBA: Mixture of Block Attention for Long-Context LLMs
von: Lu, Enzhe, et al.
Veröffentlicht: (2025)
von: Lu, Enzhe, et al.
Veröffentlicht: (2025)
Linear Attention Sequence Parallelism
von: Sun, Weigao, et al.
Veröffentlicht: (2024)
von: Sun, Weigao, et al.
Veröffentlicht: (2024)
MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts
von: Pióro, Maciej, et al.
Veröffentlicht: (2024)
von: Pióro, Maciej, et al.
Veröffentlicht: (2024)
MobileMoE: Scaling On-Device Mixture of Experts
von: Chen, Yanbei, et al.
Veröffentlicht: (2026)
von: Chen, Yanbei, et al.
Veröffentlicht: (2026)
Phase-Associative Memory: Sequence Modeling in Complex Hilbert Space
von: Vishwakarma, Gowrav, et al.
Veröffentlicht: (2026)
von: Vishwakarma, Gowrav, et al.
Veröffentlicht: (2026)
Mixture of Chapters: Scaling Learnt Memory in Transformers
von: Tibrewal, Tasmay Pankaj, et al.
Veröffentlicht: (2026)
von: Tibrewal, Tasmay Pankaj, et al.
Veröffentlicht: (2026)
MoR: Mixture of Ranks for Low-Rank Adaptation Tuning
von: Tang, Chuanyu, et al.
Veröffentlicht: (2024)
von: Tang, Chuanyu, et al.
Veröffentlicht: (2024)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
von: Lan, Michael, et al.
Veröffentlicht: (2023)
von: Lan, Michael, et al.
Veröffentlicht: (2023)
Linear Model Merging Unlocks Simple and Scalable Multimodal Data Mixture Optimization
von: Berasi, Davide, et al.
Veröffentlicht: (2026)
von: Berasi, Davide, et al.
Veröffentlicht: (2026)
MoVE: Mixture of Value Embeddings -- A New Axis for Scaling Parametric Memory in Autoregressive Models
von: Li, Yangyan
Veröffentlicht: (2026)
von: Li, Yangyan
Veröffentlicht: (2026)
MoECollab: Democratizing LLM Development Through Collaborative Mixture of Experts
von: Harshit
Veröffentlicht: (2025)
von: Harshit
Veröffentlicht: (2025)
MoESD: Mixture of Experts Stable Diffusion to Mitigate Gender Bias
von: Wang, Guorun, et al.
Veröffentlicht: (2024)
von: Wang, Guorun, et al.
Veröffentlicht: (2024)
$μ$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts
von: Koike-Akino, Toshiaki, et al.
Veröffentlicht: (2025)
von: Koike-Akino, Toshiaki, et al.
Veröffentlicht: (2025)
MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router
von: Xie, Yanyue, et al.
Veröffentlicht: (2024)
von: Xie, Yanyue, et al.
Veröffentlicht: (2024)
MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM
von: Wang, Dongwei, et al.
Veröffentlicht: (2026)
von: Wang, Dongwei, et al.
Veröffentlicht: (2026)
Assessing Episodic Memory in LLMs with Sequence Order Recall Tasks
von: Pink, Mathis, et al.
Veröffentlicht: (2024)
von: Pink, Mathis, et al.
Veröffentlicht: (2024)
Large Language Model Cascades with Mixture of Thoughts Representations for Cost-efficient Reasoning
von: Yue, Murong, et al.
Veröffentlicht: (2023)
von: Yue, Murong, et al.
Veröffentlicht: (2023)
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
von: Liu, Baihui, et al.
Veröffentlicht: (2026)
von: Liu, Baihui, et al.
Veröffentlicht: (2026)
MoSLD: An Extremely Parameter-Efficient Mixture-of-Shared LoRAs for Multi-Task Learning
von: Zhao, Lulu, et al.
Veröffentlicht: (2024)
von: Zhao, Lulu, et al.
Veröffentlicht: (2024)
MoST: Mixing Speech and Text with Modality-Aware Mixture of Experts
von: Lou, Yuxuan, et al.
Veröffentlicht: (2026)
von: Lou, Yuxuan, et al.
Veröffentlicht: (2026)
KUET at StanceNakba Shared Task: StanceMoE: Mixture-of-Experts Architecture for Stance Detection
von: Shafi, Abdullah Al, et al.
Veröffentlicht: (2026)
von: Shafi, Abdullah Al, et al.
Veröffentlicht: (2026)
LocMoE: A Low-Overhead MoE for Large Language Model Training
von: Li, Jing, et al.
Veröffentlicht: (2024)
von: Li, Jing, et al.
Veröffentlicht: (2024)
MoMQ: Mixture-of-Experts Enhances Multi-Dialect Query Generation across Relational and Non-Relational Databases
von: Lin, Zhisheng, et al.
Veröffentlicht: (2024)
von: Lin, Zhisheng, et al.
Veröffentlicht: (2024)
Towards a Comprehensive Scaling Law of Mixture-of-Experts
von: Zhao, Guoliang, et al.
Veröffentlicht: (2025)
von: Zhao, Guoliang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Liger: Linearizing Large Language Models to Gated Recurrent Structures
von: Lan, Disen, et al.
Veröffentlicht: (2025) -
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
von: Sun, Weigao, et al.
Veröffentlicht: (2025) -
Native Hybrid Attention for Efficient Sequence Modeling
von: Du, Jusen, et al.
Veröffentlicht: (2025) -
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid
von: Sun, Weigao, et al.
Veröffentlicht: (2025) -
Comba: Improving Bilinear RNNs with Closed-loop Control
von: Hu, Jiaxi, et al.
Veröffentlicht: (2025)