MoE$^2$: Optimizing Collaborative Inference for Edge Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Lyudong, Zhang, Yanning, Li, Yanhan, Wang, Shurong, Yang, Howard H., Wu, Jian, Zhang, Meng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917893809111040
author Jin, Lyudong
Zhang, Yanning
Li, Yanhan
Wang, Shurong
Yang, Howard H.
Wu, Jian
Zhang, Meng
author_facet Jin, Lyudong
Zhang, Yanning
Li, Yanhan
Wang, Shurong
Yang, Howard H.
Wu, Jian
Zhang, Meng
contents Large language models (LLMs) have demonstrated remarkable capabilities across a wide range of natural language processing tasks. Exploiting the heterogeneous capabilities of edge LLMs is crucial for diverse emerging applications, as it enables greater cost-effectiveness and reduced latency. In this work, we introduce \textit{Mixture-of-Edge-Experts (MoE$^2$)}, a novel collaborative inference framework for edge LLMs. We formulate the joint gating and expert selection problem to optimize inference performance under energy and latency constraints. Unlike conventional MoE problems, LLM expert selection is significantly more challenging due to the combinatorial nature and the heterogeneity of edge LLMs across various attributes. To this end, we propose a two-level expert selection mechanism through which we uncover an optimality-preserving property of gating parameters across expert selections. This property enables the decomposition of the training and selection processes, significantly reducing complexity. Furthermore, we leverage the objective's monotonicity and design a discrete monotonic optimization algorithm for optimal expert selection. We implement edge servers with NVIDIA Jetson AGX Orins and NVIDIA RTX 4090 GPUs, and perform extensive experiments. Our results validate that performance improvements of various LLM models and show that our MoE$^2$ method can achieve optimal trade-offs among different delay and energy budgets, and outperforms baselines under various system resource constraints.
format Preprint
id arxiv_https___arxiv_org_abs_2501_09410
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MoE$^2$: Optimizing Collaborative Inference for Edge Large Language Models
Jin, Lyudong
Zhang, Yanning
Li, Yanhan
Wang, Shurong
Yang, Howard H.
Wu, Jian
Zhang, Meng
Networking and Internet Architecture
Artificial Intelligence
Machine Learning
Large language models (LLMs) have demonstrated remarkable capabilities across a wide range of natural language processing tasks. Exploiting the heterogeneous capabilities of edge LLMs is crucial for diverse emerging applications, as it enables greater cost-effectiveness and reduced latency. In this work, we introduce \textit{Mixture-of-Edge-Experts (MoE$^2$)}, a novel collaborative inference framework for edge LLMs. We formulate the joint gating and expert selection problem to optimize inference performance under energy and latency constraints. Unlike conventional MoE problems, LLM expert selection is significantly more challenging due to the combinatorial nature and the heterogeneity of edge LLMs across various attributes. To this end, we propose a two-level expert selection mechanism through which we uncover an optimality-preserving property of gating parameters across expert selections. This property enables the decomposition of the training and selection processes, significantly reducing complexity. Furthermore, we leverage the objective's monotonicity and design a discrete monotonic optimization algorithm for optimal expert selection. We implement edge servers with NVIDIA Jetson AGX Orins and NVIDIA RTX 4090 GPUs, and perform extensive experiments. Our results validate that performance improvements of various LLM models and show that our MoE$^2$ method can achieve optimal trade-offs among different delay and energy budgets, and outperforms baselines under various system resource constraints.
title MoE$^2$: Optimizing Collaborative Inference for Edge Large Language Models
topic Networking and Internet Architecture
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2501.09410