TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications
Fuente:
arXiv
Saved in:
| Main Authors: | Bian, Zhuohang, Wu, Feiyang, Li, Zhuoran, Ma, Teng, Zhuo, Youwei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management
by: Xiong, Yi, et al.
Published: (2024)
by: Xiong, Yi, et al.
Published: (2024)
How Machine Learning-Data Driven Replication Strategies Enhance Fault Tolerance in Large-Scale Distributed Systems
by: Murimi, Almond Kiruthu
Published: (2025)
by: Murimi, Almond Kiruthu
Published: (2025)
TokenDance: Scaling Multi-Agent LLM Serving via Collective KV Cache Sharing
by: Bian, Zhuohang, et al.
Published: (2026)
by: Bian, Zhuohang, et al.
Published: (2026)
Token Coherence: Adapting MESI Cache Protocols to Minimize Synchronization Overhead in Multi-Agent LLM Systems
by: Parakhin, Vladyslav
Published: (2026)
by: Parakhin, Vladyslav
Published: (2026)
SparkAttention: High-Performance Multi-Head Attention for Large Models on Volta GPU Architecture
by: Xu, Youxuan, et al.
Published: (2025)
by: Xu, Youxuan, et al.
Published: (2025)
AutoDDL: Automatic Distributed Deep Learning with Near-Optimal Bandwidth Cost
by: Chen, Jinfan, et al.
Published: (2023)
by: Chen, Jinfan, et al.
Published: (2023)
AMP4EC: Adaptive Model Partitioning Framework for Efficient Deep Learning Inference in Edge Computing Environments
by: Zhang, Guilin, et al.
Published: (2025)
by: Zhang, Guilin, et al.
Published: (2025)
Scheduling the Unschedulable: Taming Black-Box LLM Inference at Scale
by: Yuan, Renzhong, et al.
Published: (2026)
by: Yuan, Renzhong, et al.
Published: (2026)
Accelerating Geo-distributed Machine Learning with Network-Aware Adaptive Tree and Auxiliary Route
by: Li, Zonghang, et al.
Published: (2024)
by: Li, Zonghang, et al.
Published: (2024)
HFedATM: Hierarchical Federated Domain Generalization via Optimal Transport and Regularized Mean Aggregation
by: Nguyen, Thinh, et al.
Published: (2025)
by: Nguyen, Thinh, et al.
Published: (2025)
Impact of Network Topology on Byzantine Resilience in Decentralized Federated Learning
by: Bhattacharya, Siddhartha, et al.
Published: (2024)
by: Bhattacharya, Siddhartha, et al.
Published: (2024)
From Detection to Recovery: Operational Analysis on LLM Pre-training with 504 GPUs
by: Kang, Daemyung, et al.
Published: (2026)
by: Kang, Daemyung, et al.
Published: (2026)
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
by: Ganjihal, Sanjeev Rao
Published: (2026)
by: Ganjihal, Sanjeev Rao
Published: (2026)
Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
by: Penke, Carolin, et al.
Published: (2025)
by: Penke, Carolin, et al.
Published: (2025)
FLEdge: Benchmarking Federated Machine Learning Applications in Edge Computing Systems
by: Woisetschläger, Herbert, et al.
Published: (2023)
by: Woisetschläger, Herbert, et al.
Published: (2023)
Adaptive GPU Resource Allocation for Multi-Agent Collaborative Reasoning in Serverless Environments
by: Zhang, Guilin, et al.
Published: (2025)
by: Zhang, Guilin, et al.
Published: (2025)
TAGC: Optimizing Gradient Communication in Distributed Transformer Training
by: Polyakov, Igor, et al.
Published: (2025)
by: Polyakov, Igor, et al.
Published: (2025)
GRAIN: Exact Graph Reconstruction from Gradients
by: Drencheva, Maria, et al.
Published: (2025)
by: Drencheva, Maria, et al.
Published: (2025)
Near-Optimal Sparse Allreduce for Distributed Deep Learning
by: Li, Shigang, et al.
Published: (2022)
by: Li, Shigang, et al.
Published: (2022)
FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor Cores
by: Shi, Jinliang, et al.
Published: (2024)
by: Shi, Jinliang, et al.
Published: (2024)
Chimera: Efficiently Training Large-Scale Neural Networks with Bidirectional Pipelines
by: Li, Shigang, et al.
Published: (2021)
by: Li, Shigang, et al.
Published: (2021)
Towards Building Private LLMs: Exploring Multi-Node Expert Parallelism on Apple Silicon for Mixture-of-Experts Large Language Model
by: Chen, Mu-Chi, et al.
Published: (2025)
by: Chen, Mu-Chi, et al.
Published: (2025)
Federated Learning Model Aggregation in Heterogenous Aerial and Space Networks
by: Dong, Fan, et al.
Published: (2023)
by: Dong, Fan, et al.
Published: (2023)
GraphBit: A Graph-based Agentic Framework for Non-Linear Agent Orchestration
by: Sarker, Yeahia, et al.
Published: (2026)
by: Sarker, Yeahia, et al.
Published: (2026)
Libra: Unleashing GPU Heterogeneity for High-Performance Sparse Matrix Multiplication
by: Shi, Jinliang, et al.
Published: (2025)
by: Shi, Jinliang, et al.
Published: (2025)
Aethon: A Reference-Based Replication Primitive for Constant-Time Instantiation of Stateful AI Agents
by: Rao, Swanand, et al.
Published: (2026)
by: Rao, Swanand, et al.
Published: (2026)
A Taxonomy and Resolution Strategy for Client-Level Disagreements in Federated Learning
by: Rosendal, Daan, et al.
Published: (2026)
by: Rosendal, Daan, et al.
Published: (2026)
CodeCRDT: Observation-Driven Coordination for Multi-Agent LLM Code Generation
by: Pugachev, Sergey
Published: (2025)
by: Pugachev, Sergey
Published: (2025)
zScore: A Universal Decentralised Reputation System for the Blockchain Economy
by: Udupi, Himanshu, et al.
Published: (2025)
by: Udupi, Himanshu, et al.
Published: (2025)
Federated Fine-Tuning of LLMs on the Very Edge: The Good, the Bad, the Ugly
by: Woisetschläger, Herbert, et al.
Published: (2023)
by: Woisetschläger, Herbert, et al.
Published: (2023)
Mobile Traffic Prediction at the Edge Through Distributed and Deep Transfer Learning
by: Petrella, Alfredo, et al.
Published: (2023)
by: Petrella, Alfredo, et al.
Published: (2023)
Bridging Generalization Gap of Heterogeneous Federated Clients Using Generative Models
by: Niu, Ziru, et al.
Published: (2025)
by: Niu, Ziru, et al.
Published: (2025)
TPI-LLM: Serving 70B-scale LLMs Efficiently on Low-resource Edge Devices
by: Li, Zonghang, et al.
Published: (2024)
by: Li, Zonghang, et al.
Published: (2024)
Connecting Large Language Model Agent to High Performance Computing Resource
by: Ma, Heng, et al.
Published: (2025)
by: Ma, Heng, et al.
Published: (2025)
Agent Identity URI Scheme: Topology-Independent Naming and Capability-Based Discovery for Multi-Agent Systems
by: Rodriguez Jr, Roland R.
Published: (2026)
by: Rodriguez Jr, Roland R.
Published: (2026)
Comparison of Autoscaling Frameworks for Containerised Machine-Learning-Applications in a Local and Cloud Environment
by: Schroeder, Christian, et al.
Published: (2023)
by: Schroeder, Christian, et al.
Published: (2023)
nvidia-pcm: A D-Bus-Driven Platform Configuration Manager for OpenBMC Environments
by: Singh, Harinder
Published: (2026)
by: Singh, Harinder
Published: (2026)
DDS: DPU-optimized Disaggregated Storage [Extended Report]
by: Zhang, Qizhen, et al.
Published: (2024)
by: Zhang, Qizhen, et al.
Published: (2024)
Federated Few-Shot Learning on Neuromorphic Hardware: An Empirical Study Across Physical Edge Nodes
by: Motta, Steven, et al.
Published: (2026)
by: Motta, Steven, et al.
Published: (2026)
Federated Learning for Anomaly Detection in Energy Consumption Data: Assessing the Vulnerability to Adversarial Attacks
by: Telila, Yohannis Kifle, et al.
Published: (2025)
by: Telila, Yohannis Kifle, et al.
Published: (2025)
Similar Items
-
LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management
by: Xiong, Yi, et al.
Published: (2024) -
How Machine Learning-Data Driven Replication Strategies Enhance Fault Tolerance in Large-Scale Distributed Systems
by: Murimi, Almond Kiruthu
Published: (2025) -
TokenDance: Scaling Multi-Agent LLM Serving via Collective KV Cache Sharing
by: Bian, Zhuohang, et al.
Published: (2026) -
Token Coherence: Adapting MESI Cache Protocols to Minimize Synchronization Overhead in Multi-Agent LLM Systems
by: Parakhin, Vladyslav
Published: (2026) -
SparkAttention: High-Performance Multi-Head Attention for Large Models on Volta GPU Architecture
by: Xu, Youxuan, et al.
Published: (2025)