DroidSpeak: KV Cache Sharing for Cross-LLM Communication and Multi-LLM Serving
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yuhan, Huang, Yuyang, Yao, Jiayi, Feng, Shaoting, Gu, Zhuohan, Du, Kuntai, Li, Hanchen, Cheng, Yihua, Jiang, Junchen, Lu, Shan, Musuvathi, Madan, Choukse, Esha |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
by: Feng, Shaoting, et al.
Published: (2025)
by: Feng, Shaoting, et al.
Published: (2025)
AdaptCache: KV Cache Native Storage Hierarchy for Low-Delay and High-Quality Language Model Serving
by: Feng, Shaoting, et al.
Published: (2025)
by: Feng, Shaoting, et al.
Published: (2025)
Towards More Economical Context-Augmented LLM Generation by Reusing Stored KV Cache
by: Li, Hanchen, et al.
Published: (2025)
by: Li, Hanchen, et al.
Published: (2025)
VeriCache: Turning Lossy KV Cache into Lossless LLM Inference
by: Yao, Jiayi, et al.
Published: (2026)
by: Yao, Jiayi, et al.
Published: (2026)
LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference
by: Liu, Yuhan, et al.
Published: (2025)
by: Liu, Yuhan, et al.
Published: (2025)
CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving
by: Liu, Yuhan, et al.
Published: (2023)
by: Liu, Yuhan, et al.
Published: (2023)
CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion
by: Yao, Jiayi, et al.
Published: (2024)
by: Yao, Jiayi, et al.
Published: (2024)
LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts
by: Gu, Zhuohan, et al.
Published: (2024)
by: Gu, Zhuohan, et al.
Published: (2024)
When KV Cache Reuse Fails in Multi-Agent Systems: Cross-Candidate Interaction is Crucial for LLM Judges
by: Liang, Sichu, et al.
Published: (2026)
by: Liang, Sichu, et al.
Published: (2026)
Eloquent: A More Robust Transmission Scheme for LLM Token Streaming
by: Li, Hanchen, et al.
Published: (2024)
by: Li, Hanchen, et al.
Published: (2024)
KVComm: Enabling Efficient LLM Communication through Selective KV Sharing
by: Shi, Xiangyu, et al.
Published: (2025)
by: Shi, Xiangyu, et al.
Published: (2025)
Pancake: Hierarchical Memory System for Multi-Agent LLM Serving
by: Hu, Zhengding, et al.
Published: (2026)
by: Hu, Zhengding, et al.
Published: (2026)
Speaking the Language of Teamwork: LLM-Guided Credit Assignment in Multi-Agent Reinforcement Learning
by: Lin, Muhan, et al.
Published: (2025)
by: Lin, Muhan, et al.
Published: (2025)
Murakkab: Resource-Efficient Agentic Workflow Orchestration in Cloud Platforms
by: Chaudhry, Gohar Irfan, et al.
Published: (2025)
by: Chaudhry, Gohar Irfan, et al.
Published: (2025)
Sherlock: Reliable and Efficient Agentic Workflow Execution
by: Ro, Yeonju, et al.
Published: (2025)
by: Ro, Yeonju, et al.
Published: (2025)
QKVShare: Quantized KV-Cache Handoff for Multi-Agent On-Device LLMs
by: Honavar, Pratik, et al.
Published: (2026)
by: Honavar, Pratik, et al.
Published: (2026)
KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent Systems
by: Ye, Hancheng, et al.
Published: (2025)
by: Ye, Hancheng, et al.
Published: (2025)
Self-Organizing Agent Network for LLM-based Workflow Automation
by: Xiong, Yiming, et al.
Published: (2025)
by: Xiong, Yiming, et al.
Published: (2025)
Q-KVComm: Efficient Multi-Agent Communication Via Adaptive KV Cache Compression
by: Kriuk, Boris, et al.
Published: (2025)
by: Kriuk, Boris, et al.
Published: (2025)
Guiding LLM-Based Human Mobility Simulation with Mobility Measures from Shared Data
by: Yan, Hua, et al.
Published: (2026)
by: Yan, Hua, et al.
Published: (2026)
Efficient LLM Serving for Agentic Workflows: A Data Systems Perspective
by: Wadlom, Noppanat, et al.
Published: (2026)
by: Wadlom, Noppanat, et al.
Published: (2026)
Pythia: Exploiting Workflow Predictability for Efficient Agent-Native LLM Serving
by: Yu, Shan, et al.
Published: (2026)
by: Yu, Shan, et al.
Published: (2026)
Evidence-Decision-Feedback: Theory-Driven Adaptive Scaffolding for LLM Agents
by: Cohn, Clayton, et al.
Published: (2026)
by: Cohn, Clayton, et al.
Published: (2026)
ShadowServe: Interference-Free KV Cache Fetching for Distributed Prefix Caching
by: Xiang, Xingyu, et al.
Published: (2025)
by: Xiang, Xingyu, et al.
Published: (2025)
A Theory-Guided LLM Pedagogical Agent for STEM+C Scaffolding Without Over-Reliance
by: Cohn, Clayton, et al.
Published: (2026)
by: Cohn, Clayton, et al.
Published: (2026)
KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows
by: Pan, Zaifeng, et al.
Published: (2025)
by: Pan, Zaifeng, et al.
Published: (2025)
LLM-Guided Strategy Synthesis for Scalable Equality Saturation
by: Yin, Chenyun, et al.
Published: (2026)
by: Yin, Chenyun, et al.
Published: (2026)
LLM-ABM for Transportation: Assessing the Potential of LLM Agents in System Analysis
by: Liu, Tianming, et al.
Published: (2025)
by: Liu, Tianming, et al.
Published: (2025)
Competition and Cooperation of LLM Agents in Games
by: Yao, Jiayi, et al.
Published: (2026)
by: Yao, Jiayi, et al.
Published: (2026)
Ensuring Fair LLM Serving Amid Diverse Applications
by: Khan, Redwan Ibne Seraj, et al.
Published: (2024)
by: Khan, Redwan Ibne Seraj, et al.
Published: (2024)
Do Large Language Models Need a Content Delivery Network?
by: Cheng, Yihua, et al.
Published: (2024)
by: Cheng, Yihua, et al.
Published: (2024)
Scaling Teams or Scaling Time? Memory Enabled Lifelong Learning in LLM Multi-Agent Systems
by: Wu, Shanglin, et al.
Published: (2026)
by: Wu, Shanglin, et al.
Published: (2026)
MARLIN: Multi-Agent Reinforcement Learning with Murmuration Intelligence and LLM Guidance for Reservoir Management
by: Fu, Heming, et al.
Published: (2025)
by: Fu, Heming, et al.
Published: (2025)
Learning to Speak on Behalf of a Group: Medium Access Control for Sending a Shared Message
by: Haque, Shaan ul, et al.
Published: (2022)
by: Haque, Shaan ul, et al.
Published: (2022)
Planner-Auditor Twin: Agentic Discharge Planning with FHIR-Based LLM Planning, Guideline Recall, Optional Caching and Self-Improvement
by: Wu, Kaiyuan, et al.
Published: (2026)
by: Wu, Kaiyuan, et al.
Published: (2026)
Negotiating Comfort: Simulating Personality-Driven LLM Agents in Shared Residential Social Networks
by: Rende, Ann Nedime Nese, et al.
Published: (2025)
by: Rende, Ann Nedime Nese, et al.
Published: (2025)
CASPIAN: Online Detection and Attribution of Cascade Attacks in LLM Multi-Agent Systems via Cross-Channel Causal Monitoring
by: Venkatesh, Kavana, et al.
Published: (2026)
by: Venkatesh, Kavana, et al.
Published: (2026)
LLM-ALSO: LLM-Driven Adaptive Learning-Signal Optimization for Multi-Agent Reinforcement Learning
by: Wu, Xiaoguang, et al.
Published: (2026)
by: Wu, Xiaoguang, et al.
Published: (2026)
MACRO-LLM: LLM-Empowered Multi-Agent Collaborative Reasoning under Spatiotemporal Partial Observability
by: Chen, Handi, et al.
Published: (2026)
by: Chen, Handi, et al.
Published: (2026)
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing
by: Amico, Jeffrey, et al.
Published: (2025)
by: Amico, Jeffrey, et al.
Published: (2025)
Similar Items
-
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
by: Feng, Shaoting, et al.
Published: (2025) -
AdaptCache: KV Cache Native Storage Hierarchy for Low-Delay and High-Quality Language Model Serving
by: Feng, Shaoting, et al.
Published: (2025) -
Towards More Economical Context-Augmented LLM Generation by Reusing Stored KV Cache
by: Li, Hanchen, et al.
Published: (2025) -
VeriCache: Turning Lossy KV Cache into Lossless LLM Inference
by: Yao, Jiayi, et al.
Published: (2026) -
LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference
by: Liu, Yuhan, et al.
Published: (2025)