MoE$^2$: Optimizing Collaborative Inference for Edge Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Jin, Lyudong, Zhang, Yanning, Li, Yanhan, Wang, Shurong, Yang, Howard H., Wu, Jian, Zhang, Meng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CoMoE: Collaborative Optimization of Expert Aggregation and Offloading for MoE-based LLMs at Edge
by: Li, Muqing, et al.
Published: (2025)
by: Li, Muqing, et al.
Published: (2025)
Asynchronous Fractional Multi-Agent Deep Reinforcement Learning for Age-Minimal Mobile Edge Computing
by: Jin, Lyudong, et al.
Published: (2024)
by: Jin, Lyudong, et al.
Published: (2024)
The MoE-Empowered Edge LLMs Deployment: Architecture, Challenges, and Opportunities
by: Li, Ning, et al.
Published: (2025)
by: Li, Ning, et al.
Published: (2025)
CE-LSLM: Efficient Large-Small Language Model Inference and Communication via Cloud-Edge Collaboration
by: Zhu, Pengyan, et al.
Published: (2025)
by: Zhu, Pengyan, et al.
Published: (2025)
Cooperative Edge Caching with Large Language Model in Wireless Networks
by: Yang, Ning, et al.
Published: (2026)
by: Yang, Ning, et al.
Published: (2026)
An Edge-Cloud Collaboration Framework for Generative AI Service Provision with Synergetic Big Cloud Model and Small Edge Models
by: Tian, Yuqing, et al.
Published: (2024)
by: Tian, Yuqing, et al.
Published: (2024)
UAV-Assisted Cooperative Edge Inference for Low-Altitude Economy via MoE-based Hierarchical Deep Reinforcement Learning
by: Zhuang, Wenhao, et al.
Published: (2026)
by: Zhuang, Wenhao, et al.
Published: (2026)
SiftMoE: Similarity-Aware Energy-Efficient Expert Selection for Wireless Distributed MoE Inference
by: Chen, Qian, et al.
Published: (2026)
by: Chen, Qian, et al.
Published: (2026)
Edge Intelligence Optimization for Large Language Model Inference with Batching and Quantization
by: Zhang, Xinyuan, et al.
Published: (2024)
by: Zhang, Xinyuan, et al.
Published: (2024)
Toward Resource-Efficient Collaboration of Large AI Models in Mobile Edge Networks
by: Li, Peichun, et al.
Published: (2026)
by: Li, Peichun, et al.
Published: (2026)
Graph Attention Reinforcement Learning for Multicast Routing and Age-Optimal Scheduling
by: Zhang, Yanning, et al.
Published: (2024)
by: Zhang, Yanning, et al.
Published: (2024)
Age of Information in Random Access Networks with Energy Harvesting
by: Zhao, Fangming, et al.
Published: (2024)
by: Zhao, Fangming, et al.
Published: (2024)
Congestion Control System Optimization with Large Language Models
by: He, Zhiyuan, et al.
Published: (2025)
by: He, Zhiyuan, et al.
Published: (2025)
Hallucination-aware Optimization for Large Language Model-empowered Communications
by: Liu, Yinqiu, et al.
Published: (2024)
by: Liu, Yinqiu, et al.
Published: (2024)
Adaptive Contextual Caching for Mobile Edge Large Language Model Service
by: Liu, Guangyuan, et al.
Published: (2025)
by: Liu, Guangyuan, et al.
Published: (2025)
Optimizing System Latency for Blockchain-Encrypted Edge Computing in Internet of Vehicles
by: Zhang, Cui, et al.
Published: (2025)
by: Zhang, Cui, et al.
Published: (2025)
Cloud-Edge Collaborative Large Models for Robust Photovoltaic Power Forecasting
by: Qiao, Nan, et al.
Published: (2026)
by: Qiao, Nan, et al.
Published: (2026)
Design Insights into Partition Placement and Routing for DNN Inference in Multi-Hop Edge Networks
by: Zhang, Jinkun, et al.
Published: (2026)
by: Zhang, Jinkun, et al.
Published: (2026)
Large Language Model-Based Task Offloading and Resource Allocation for Digital Twin Edge Computing Networks
by: Wu, Qiong, et al.
Published: (2025)
by: Wu, Qiong, et al.
Published: (2025)
SLIDE: Simultaneous Model Downloading and Inference at the Wireless Network Edge
by: Qu, Guanqiao, et al.
Published: (2025)
by: Qu, Guanqiao, et al.
Published: (2025)
A Confidence-Constrained Cloud-Edge Collaborative Framework for Autism Spectrum Disorder Diagnosis
by: Deng, Qi, et al.
Published: (2025)
by: Deng, Qi, et al.
Published: (2025)
PIB: Prioritized Information Bottleneck Framework for Collaborative Edge Video Analytics
by: Fang, Zhengru, et al.
Published: (2024)
by: Fang, Zhengru, et al.
Published: (2024)
Task-oriented Age of Information for Remote Inference with Hybrid Language Models
by: Gan, Shuying, et al.
Published: (2025)
by: Gan, Shuying, et al.
Published: (2025)
CREWS: Collaborative Robust Edge WiFi Sensing with Asynchronous and Incomplete Observations
by: Chen, Yinan, et al.
Published: (2026)
by: Chen, Yinan, et al.
Published: (2026)
Chain-of-Thought for Large Language Model-empowered Wireless Communications
by: Wang, Xudong, et al.
Published: (2025)
by: Wang, Xudong, et al.
Published: (2025)
Hybrid RIS-Aided Digital Over-the-Air Computing for Edge AI Inference: Joint Feature Quantization and Active-Passive Beamforming Design
by: Fu, Yang, et al.
Published: (2025)
by: Fu, Yang, et al.
Published: (2025)
SpaceMoE: Towards Orbital General Intelligence with Distributed Mixture-of-Experts Inference
by: Chen, Qian, et al.
Published: (2026)
by: Chen, Qian, et al.
Published: (2026)
Large Models for Aerial Edges: An Edge-Cloud Model Evolution and Communication Paradigm
by: Zhang, Shuhang, et al.
Published: (2024)
by: Zhang, Shuhang, et al.
Published: (2024)
Age-of-Information and Energy Optimization in Digital Twin Edge Networks
by: Guo, Yongna, et al.
Published: (2024)
by: Guo, Yongna, et al.
Published: (2024)
On the Optimization of Model Aggregation for Federated Learning at the Network Edge
by: Li, Mengyao, et al.
Published: (2025)
by: Li, Mengyao, et al.
Published: (2025)
Diffusion Model-based Incentive Mechanism with Prospect Theory for Edge AIGC Services in 6G IoT
by: Wen, Jinbo, et al.
Published: (2024)
by: Wen, Jinbo, et al.
Published: (2024)
Constraint-Compliant Network Optimization through Large Language Models
by: Song, Youngjin, et al.
Published: (2025)
by: Song, Youngjin, et al.
Published: (2025)
Optimizing Multi-Gateway LoRaWAN via Cloud-Edge Collaboration and Knowledge Distillation
by: Yang, Hong
Published: (2025)
by: Yang, Hong
Published: (2025)
Energy-Efficient Federated Learning and Migration in Digital Twin Edge Networks
by: Zhou, Yuzhi, et al.
Published: (2025)
by: Zhou, Yuzhi, et al.
Published: (2025)
RRTO: A High-Performance Transparent Offloading System for Model Inference in Mobile Edge Computing
by: Sun, Zekai, et al.
Published: (2025)
by: Sun, Zekai, et al.
Published: (2025)
FluxShard: Motion-Aware Feature Cache Reuse for Collaborative Video Analytics in Mobile Edge Computing
by: Guan, Xiuxian, et al.
Published: (2026)
by: Guan, Xiuxian, et al.
Published: (2026)
On Network-Aware Semantic Communication and Edge-Cloud Collaborative Intelligence Systems
by: Nasif, Murdadha, et al.
Published: (2025)
by: Nasif, Murdadha, et al.
Published: (2025)
A Paradigm For Collaborative Pervasive Fog Computing Ecosystems at the Network Edge
by: Mtibaa, Abderrahmen
Published: (2024)
by: Mtibaa, Abderrahmen
Published: (2024)
Cluster Topology-Driven Placement of Experts Reduces Network Traffic in MoE Inference
by: Sivtsov, Danil, et al.
Published: (2025)
by: Sivtsov, Danil, et al.
Published: (2025)
LLM-Slice: Dedicated Wireless Network Slicing for Large Language Models
by: Liu, Boyi, et al.
Published: (2024)
by: Liu, Boyi, et al.
Published: (2024)
Similar Items
-
CoMoE: Collaborative Optimization of Expert Aggregation and Offloading for MoE-based LLMs at Edge
by: Li, Muqing, et al.
Published: (2025) -
Asynchronous Fractional Multi-Agent Deep Reinforcement Learning for Age-Minimal Mobile Edge Computing
by: Jin, Lyudong, et al.
Published: (2024) -
The MoE-Empowered Edge LLMs Deployment: Architecture, Challenges, and Opportunities
by: Li, Ning, et al.
Published: (2025) -
CE-LSLM: Efficient Large-Small Language Model Inference and Communication via Cloud-Edge Collaboration
by: Zhu, Pengyan, et al.
Published: (2025) -
Cooperative Edge Caching with Large Language Model in Wireless Networks
by: Yang, Ning, et al.
Published: (2026)