ShardMemo: Masked MoE Routing for Sharded Agentic LLM Memory
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Yang, Dai, Chengxiao, Xiu, Yue, Kou, Mengying, Zheng, Yuliang, Niyato, Dusit |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MEMOREPAIR: Barrier-First Cascade Repair in Agentic Memory
by: Zhao, Yang, et al.
Published: (2026)
by: Zhao, Yang, et al.
Published: (2026)
CLAUSE: Agentic Neuro-Symbolic Knowledge Graph Reasoning via Dynamic Learnable Context Engineering
by: Zhao, Yang, et al.
Published: (2025)
by: Zhao, Yang, et al.
Published: (2025)
From Graph Retrieval to Schema Realization: Counterfactual Validation for Text-to-SPARQL over Heterogeneous Knowledge Graphs
by: Zhao, Yang, et al.
Published: (2025)
by: Zhao, Yang, et al.
Published: (2025)
Accelerating MoE Model Inference with Expert Sharding
by: Balmau, Oana, et al.
Published: (2025)
by: Balmau, Oana, et al.
Published: (2025)
Contextual Agentic Memory is a Memo, Not True Memory
by: Xu, Binyan, et al.
Published: (2026)
by: Xu, Binyan, et al.
Published: (2026)
MemoBrain: Executive Memory as an Agentic Brain for Reasoning
by: Qian, Hongjin, et al.
Published: (2026)
by: Qian, Hongjin, et al.
Published: (2026)
BEAM: Binary Expert Activation Masking for Dynamic Routing in MoE
by: Wu, Juntong, et al.
Published: (2026)
by: Wu, Juntong, et al.
Published: (2026)
Model Agnostic Hybrid Sharding For Heterogeneous Distributed Inference
by: Angione, Claudio, et al.
Published: (2024)
by: Angione, Claudio, et al.
Published: (2024)
Scalable Federated Unlearning via Isolated and Coded Sharding
by: Lin, Yijing, et al.
Published: (2024)
by: Lin, Yijing, et al.
Published: (2024)
ArcMemo: Abstract Reasoning Composition with Lifelong LLM Memory
by: Ho, Matthew, et al.
Published: (2025)
by: Ho, Matthew, et al.
Published: (2025)
MiSS: Revisiting the Trade-off in LoRA with an Efficient Shard-Sharing Structure
by: Kang, Jiale, et al.
Published: (2024)
by: Kang, Jiale, et al.
Published: (2024)
LLaDA-MoE: A Sparse MoE Diffusion Language Model
by: Zhu, Fengqi, et al.
Published: (2025)
by: Zhu, Fengqi, et al.
Published: (2025)
Unveiling Language Routing Isolation in Multilingual MoE Models for Interpretable Subnetwork Adaptation
by: Zheng, Kening, et al.
Published: (2026)
by: Zheng, Kening, et al.
Published: (2026)
Movable Antenna Enhanced Federated Fine-Tuning of Large Language Models via Hybrid Client Selection Optimization
by: Zhao, Yang, et al.
Published: (2025)
by: Zhao, Yang, et al.
Published: (2025)
Polysemantic Experts, Monosemantic Paths: Routing as Control in MoEs
by: Ye, Charles, et al.
Published: (2026)
by: Ye, Charles, et al.
Published: (2026)
MemoRAG: Boosting Long Context Processing with Global Memory-Enhanced Retrieval Augmentation
by: Qian, Hongjin, et al.
Published: (2024)
by: Qian, Hongjin, et al.
Published: (2024)
Dynamic Language Group-Based MoE: Enhancing Code-Switching Speech Recognition with Hierarchical Routing
by: Huang, Hukai, et al.
Published: (2024)
by: Huang, Hukai, et al.
Published: (2024)
Sigma-MoE-Tiny Technical Report
by: Hu, Qingguo, et al.
Published: (2025)
by: Hu, Qingguo, et al.
Published: (2025)
Evaluating Expert Contributions in a MoE LLM for Quiz-Based Tasks
by: Chernov, Andrei
Published: (2025)
by: Chernov, Andrei
Published: (2025)
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
by: Ludziejewski, Jan, et al.
Published: (2025)
by: Ludziejewski, Jan, et al.
Published: (2025)
LocMoE: A Low-Overhead MoE for Large Language Model Training
by: Li, Jing, et al.
Published: (2024)
by: Li, Jing, et al.
Published: (2024)
MoRAL: MoE Augmented LoRA for LLMs' Lifelong Learning
by: Yang, Shu, et al.
Published: (2024)
by: Yang, Shu, et al.
Published: (2024)
Mix-MoE: Improving Multilingual Machine Translation of Large Language Models through Mixed MoEs
by: Li, Bo, et al.
Published: (2026)
by: Li, Bo, et al.
Published: (2026)
AquilaMoE: Efficient Training for MoE Models with Scale-Up and Scale-Out Strategies
by: Zhang, Bo-Wen, et al.
Published: (2024)
by: Zhang, Bo-Wen, et al.
Published: (2024)
Yuan3.0 Ultra: A Trillion-Parameter Enterprise-Oriented MoE LLM
by: ai, YuanLab., et al.
Published: (2026)
by: ai, YuanLab., et al.
Published: (2026)
OmniMoE: An Efficient MoE by Orchestrating Atomic Experts at Scale
by: Shi, Jingze, et al.
Published: (2026)
by: Shi, Jingze, et al.
Published: (2026)
GRIN: GRadient-INformed MoE
by: Liu, Liyuan, et al.
Published: (2024)
by: Liu, Liyuan, et al.
Published: (2024)
Ada-K Routing: Boosting the Efficiency of MoE-based LLMs
by: Yue, Tongtian, et al.
Published: (2024)
by: Yue, Tongtian, et al.
Published: (2024)
SP-Chain: Boosting Intra-Shard and Cross-Shard Security and Performance in Blockchain Sharding
by: Li, Mingzhe, et al.
Published: (2024)
by: Li, Mingzhe, et al.
Published: (2024)
Efficient Prompting for LLM-based Generative Internet of Things
by: Xiao, Bin, et al.
Published: (2024)
by: Xiao, Bin, et al.
Published: (2024)
MoE-SpAc: Efficient MoE Inference Based on Speculative Activation Utility in Heterogeneous Edge Scenarios
by: Li, Shuhuai, et al.
Published: (2026)
by: Li, Shuhuai, et al.
Published: (2026)
S2-Attention: Hardware-Aware Context Sharding Among Attention Heads
by: Lin, Xihui, et al.
Published: (2024)
by: Lin, Xihui, et al.
Published: (2024)
SMAR: Soft Modality-Aware Routing Strategy for MoE-based Multimodal Large Language Models Preserving Language Capabilities
by: Xia, Guoyang, et al.
Published: (2025)
by: Xia, Guoyang, et al.
Published: (2025)
MoE-Sieve: Routing-Guided LoRA for Efficient MoE Fine-Tuning
by: Manzoni, Andrea
Published: (2026)
by: Manzoni, Andrea
Published: (2026)
Sharded Elimination and Combining for Highly-Efficient Concurrent Stacks
by: Singh, Ajay, et al.
Published: (2026)
by: Singh, Ajay, et al.
Published: (2026)
Towards Specialized Generalists: A Multi-Task MoE-LoRA Framework for Domain-Specific LLM Adaptation
by: Yang, Yuxin, et al.
Published: (2026)
by: Yang, Yuxin, et al.
Published: (2026)
ACCO: Accumulate While You Communicate for Communication-Overlapped Sharded LLM Training
by: Nabli, Adel, et al.
Published: (2024)
by: Nabli, Adel, et al.
Published: (2024)
DynaShard: Secure and Adaptive Blockchain Sharding Protocol with Hybrid Consensus and Dynamic Shard Management
by: Liu, Ao, et al.
Published: (2024)
by: Liu, Ao, et al.
Published: (2024)
Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding
by: Bhatia, Nidhi, et al.
Published: (2025)
by: Bhatia, Nidhi, et al.
Published: (2025)
Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers
by: Ma, Wenhan, et al.
Published: (2025)
by: Ma, Wenhan, et al.
Published: (2025)
Similar Items
-
MEMOREPAIR: Barrier-First Cascade Repair in Agentic Memory
by: Zhao, Yang, et al.
Published: (2026) -
CLAUSE: Agentic Neuro-Symbolic Knowledge Graph Reasoning via Dynamic Learnable Context Engineering
by: Zhao, Yang, et al.
Published: (2025) -
From Graph Retrieval to Schema Realization: Counterfactual Validation for Text-to-SPARQL over Heterogeneous Knowledge Graphs
by: Zhao, Yang, et al.
Published: (2025) -
Accelerating MoE Model Inference with Expert Sharding
by: Balmau, Oana, et al.
Published: (2025) -
Contextual Agentic Memory is a Memo, Not True Memory
by: Xu, Binyan, et al.
Published: (2026)