Splitwise: Collaborative Edge-Cloud Inference for LLMs via Lyapunov-Assisted DRL

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Younesi, Abolfazl, Maryan, Abbas Shabrang, Oustad, Elyas, Samani, Zahra Najafabadi, Ansari, Mohsen, Fahringer, Thomas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908737216708608
author Younesi, Abolfazl
Maryan, Abbas Shabrang
Oustad, Elyas
Samani, Zahra Najafabadi
Ansari, Mohsen
Fahringer, Thomas
author_facet Younesi, Abolfazl
Maryan, Abbas Shabrang
Oustad, Elyas
Samani, Zahra Najafabadi
Ansari, Mohsen
Fahringer, Thomas
contents Deploying large language models (LLMs) on edge devices is challenging due to their limited memory and power resources. Cloud-only inference reduces device burden but introduces high latency and cost. Static edge-cloud partitions optimize a single metric and struggle when bandwidth fluctuates. We propose Splitwise, a novel Lyapunov-assisted deep reinforcement learning (DRL) framework for fine-grained, adaptive partitioning of LLMs across edge and cloud environments. Splitwise decomposes transformer layers into attention heads and feed-forward sub-blocks, exposing more partition choices than layer-wise schemes. A hierarchical DRL policy, guided by Lyapunov optimization, jointly minimizes latency, energy consumption, and accuracy degradation while guaranteeing queue stability under stochastic workloads and variable network bandwidth. Splitwise also guarantees robustness via partition checkpoints with exponential backoff recovery in case of communication failures. Experiments on Jetson Orin NX, Galaxy S23, and Raspberry Pi 5 with GPT-2 (1.5B), LLaMA-7B, and LLaMA-13B show that Splitwise reduces end-to-end latency by 1.4x-2.8x and cuts energy consumption by up to 41% compared with existing partitioners. It lowers the 95th-percentile latency by 53-61% relative to cloud-only execution, while maintaining accuracy and modest memory requirements.
format Preprint
id arxiv_https___arxiv_org_abs_2512_23310
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Splitwise: Collaborative Edge-Cloud Inference for LLMs via Lyapunov-Assisted DRL
Younesi, Abolfazl
Maryan, Abbas Shabrang
Oustad, Elyas
Samani, Zahra Najafabadi
Ansari, Mohsen
Fahringer, Thomas
Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Emerging Technologies
Networking and Internet Architecture
68M14, 68T05, 68T07, 90C40
I.2.6; I.2.11; C.2.4; C.5.3
Deploying large language models (LLMs) on edge devices is challenging due to their limited memory and power resources. Cloud-only inference reduces device burden but introduces high latency and cost. Static edge-cloud partitions optimize a single metric and struggle when bandwidth fluctuates. We propose Splitwise, a novel Lyapunov-assisted deep reinforcement learning (DRL) framework for fine-grained, adaptive partitioning of LLMs across edge and cloud environments. Splitwise decomposes transformer layers into attention heads and feed-forward sub-blocks, exposing more partition choices than layer-wise schemes. A hierarchical DRL policy, guided by Lyapunov optimization, jointly minimizes latency, energy consumption, and accuracy degradation while guaranteeing queue stability under stochastic workloads and variable network bandwidth. Splitwise also guarantees robustness via partition checkpoints with exponential backoff recovery in case of communication failures. Experiments on Jetson Orin NX, Galaxy S23, and Raspberry Pi 5 with GPT-2 (1.5B), LLaMA-7B, and LLaMA-13B show that Splitwise reduces end-to-end latency by 1.4x-2.8x and cuts energy consumption by up to 41% compared with existing partitioners. It lowers the 95th-percentile latency by 53-61% relative to cloud-only execution, while maintaining accuracy and modest memory requirements.
title Splitwise: Collaborative Edge-Cloud Inference for LLMs via Lyapunov-Assisted DRL
topic Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Emerging Technologies
Networking and Internet Architecture
68M14, 68T05, 68T07, 90C40
I.2.6; I.2.11; C.2.4; C.5.3
url https://arxiv.org/abs/2512.23310