DisCEdge: Distributed Context Management for Large Language Models at the Edge

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Malekabbasi, Mohammadreza, Wang, Minghe, Bermbach, David
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915922759909376
author Malekabbasi, Mohammadreza
Wang, Minghe
Bermbach, David
author_facet Malekabbasi, Mohammadreza
Wang, Minghe
Bermbach, David
contents Deploying Large Language Model (LLM) services at the edge benefits latency-sensitive and privacy-aware applications. However, the stateless nature of LLMs makes managing user context (e.g., sessions, preferences) across geo-distributed edge nodes challenging. Existing solutions, such as client-side context storage, introduce network latency and bandwidth overhead, undermining edge deployment advantages. We propose DisCEdge, a distributed context management system that stores and replicates user context in tokenized form across edge nodes. By maintaining context as token sequences, our system avoids redundant computation and enables efficient data replication. We evaluate an open-source prototype in a realistic edge environment. DisCEdge improves median response times by up to 14.46% and lowers median inter-node synchronization overhead by up to 15% compared to a raw-text-based system. It also reduces client request sizes by a median of 90% compared to client-side context management, while guaranteeing data consistency.
format Preprint
id arxiv_https___arxiv_org_abs_2511_22599
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DisCEdge: Distributed Context Management for Large Language Models at the Edge
Malekabbasi, Mohammadreza
Wang, Minghe
Bermbach, David
Distributed, Parallel, and Cluster Computing
Databases
Machine Learning
Deploying Large Language Model (LLM) services at the edge benefits latency-sensitive and privacy-aware applications. However, the stateless nature of LLMs makes managing user context (e.g., sessions, preferences) across geo-distributed edge nodes challenging. Existing solutions, such as client-side context storage, introduce network latency and bandwidth overhead, undermining edge deployment advantages. We propose DisCEdge, a distributed context management system that stores and replicates user context in tokenized form across edge nodes. By maintaining context as token sequences, our system avoids redundant computation and enables efficient data replication. We evaluate an open-source prototype in a realistic edge environment. DisCEdge improves median response times by up to 14.46% and lowers median inter-node synchronization overhead by up to 15% compared to a raw-text-based system. It also reduces client request sizes by a median of 90% compared to client-side context management, while guaranteeing data consistency.
title DisCEdge: Distributed Context Management for Large Language Models at the Edge
topic Distributed, Parallel, and Cluster Computing
Databases
Machine Learning
url https://arxiv.org/abs/2511.22599