KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows
Fuente:
arXiv
Salvato in:
| Autori principali: | Pan, Zaifeng, Patel, Ajjkumar, Hu, Zhengding, Shen, Yipeng, Guan, Yue, Li, Wan-Lu, Qin, Lianhui, Wang, Yida, Ding, Yufei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ScaleSim: Serving Large-Scale Multi-Agent Simulation with Invocation Distance-Based Memory Management
di: Pan, Zaifeng, et al.
Pubblicazione: (2026)
di: Pan, Zaifeng, et al.
Pubblicazione: (2026)
Empowering Scientific Workflows with Federated Agents
di: Kamatar, Alok, et al.
Pubblicazione: (2025)
di: Kamatar, Alok, et al.
Pubblicazione: (2025)
FlashEvolve: Accelerating Agent Self-Evolution with Asynchronous Stage Orchestration
di: Hu, Zhengding, et al.
Pubblicazione: (2026)
di: Hu, Zhengding, et al.
Pubblicazione: (2026)
Aegis: Taxonomy and Optimizations for Overcoming Agent-Environment Failures in LLM Agents
di: Song, Kevin, et al.
Pubblicazione: (2025)
di: Song, Kevin, et al.
Pubblicazione: (2025)
Pythia: Exploiting Workflow Predictability for Efficient Agent-Native LLM Serving
di: Yu, Shan, et al.
Pubblicazione: (2026)
di: Yu, Shan, et al.
Pubblicazione: (2026)
RcLLM: Accelerating Generative Recommendation via Beyond-Prefix KV Caching
di: Zhao, Zhan, et al.
Pubblicazione: (2026)
di: Zhao, Zhan, et al.
Pubblicazione: (2026)
UFO3: Weaving the Digital Agent Galaxy
di: Zhang, Chaoyun, et al.
Pubblicazione: (2025)
di: Zhang, Chaoyun, et al.
Pubblicazione: (2025)
Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap
di: Qiang, Xinwei, et al.
Pubblicazione: (2026)
di: Qiang, Xinwei, et al.
Pubblicazione: (2026)
Acceleration of Gossip Algorithms through the Euler-Poisson-Darboux Equation
di: Berthier, Raphaël, et al.
Pubblicazione: (2022)
di: Berthier, Raphaël, et al.
Pubblicazione: (2022)
DynTaskMAS: A Dynamic Task Graph-driven Framework for Asynchronous and Parallel LLM-based Multi-Agent Systems
di: Yu, Junwei, et al.
Pubblicazione: (2025)
di: Yu, Junwei, et al.
Pubblicazione: (2025)
On the Limits of Information Spread by Memory-less Agents
di: D'Archivio, Niccolò, et al.
Pubblicazione: (2024)
di: D'Archivio, Niccolò, et al.
Pubblicazione: (2024)
Near-linear Time Dispersion of Mobile Agents
di: Sudo, Yuichi, et al.
Pubblicazione: (2023)
di: Sudo, Yuichi, et al.
Pubblicazione: (2023)
Distributed Butterfly Analysis using Mobile Agents
di: Chand, Prabhat Kumar, et al.
Pubblicazione: (2025)
di: Chand, Prabhat Kumar, et al.
Pubblicazione: (2025)
Computing Tree Structures in Anonymous Graphs via Mobile Agents
di: Chand, Prabhat Kumar, et al.
Pubblicazione: (2025)
di: Chand, Prabhat Kumar, et al.
Pubblicazione: (2025)
Distributed Set-membership Filtering Frameworks For Multi-agent Systems With Absolute and Relative Measurements
di: Ding, Yu, et al.
Pubblicazione: (2023)
di: Ding, Yu, et al.
Pubblicazione: (2023)
Decoupling Correctness from Policy: A Deterministic Causal Structure for Multi-Agent Systems
di: Ren, Zhiyuan, et al.
Pubblicazione: (2025)
di: Ren, Zhiyuan, et al.
Pubblicazione: (2025)
Agent-Based Triangle Counting: Unlocking Truss Decomposition, Triangle Centrality, and Local Clustering Coefficient
di: Chand, Prabhat Kumar, et al.
Pubblicazione: (2024)
di: Chand, Prabhat Kumar, et al.
Pubblicazione: (2024)
When Agents Control Robots: A Zero Trust Policy Model for Agentic Cyber-Physical Systems
di: Ranathunga, Tharindu, et al.
Pubblicazione: (2026)
di: Ranathunga, Tharindu, et al.
Pubblicazione: (2026)
APWA: A Distributed Architecture for Parallelizable Agentic Workflows
di: Rose, Evan, et al.
Pubblicazione: (2026)
di: Rose, Evan, et al.
Pubblicazione: (2026)
Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic
di: Liu, Shuo, et al.
Pubblicazione: (2026)
di: Liu, Shuo, et al.
Pubblicazione: (2026)
VineLM: Trie-Based Fine-Grained Control for Agentic Workflows
di: Pagonas, Nikos, et al.
Pubblicazione: (2026)
di: Pagonas, Nikos, et al.
Pubblicazione: (2026)
Optimizing PyTorch Inference with LLM-Based Multi-Agent Systems
di: Nagaitsev, Kirill, et al.
Pubblicazione: (2025)
di: Nagaitsev, Kirill, et al.
Pubblicazione: (2025)
Marconi: Prefix Caching for the Era of Hybrid LLMs
di: Pan, Rui, et al.
Pubblicazione: (2024)
di: Pan, Rui, et al.
Pubblicazione: (2024)
Towards Blockchain-based Multi-Agent Robotic Systems: Analysis, Classification and Applications
di: Afanasyev, Ilya, et al.
Pubblicazione: (2019)
di: Afanasyev, Ilya, et al.
Pubblicazione: (2019)
An Accelerated Distributed Stochastic Gradient Method with Momentum
di: Huang, Kun, et al.
Pubblicazione: (2024)
di: Huang, Kun, et al.
Pubblicazione: (2024)
Agentic Fog: A Policy-driven Framework for Distributed Intelligence in Fog Computing
di: Akbar, Saeed, et al.
Pubblicazione: (2026)
di: Akbar, Saeed, et al.
Pubblicazione: (2026)
Ledger-State Stigmergy: A Formal Framework for Indirect Coordination Grounded in Distributed Ledger State
di: García, Fernando Paredes
Pubblicazione: (2026)
di: García, Fernando Paredes
Pubblicazione: (2026)
When Computing follows Vehicles: Decentralized Mobility-Aware Resource Allocation for Edge-to-Cloud Continuum
di: Nezami, Zeinab, et al.
Pubblicazione: (2024)
di: Nezami, Zeinab, et al.
Pubblicazione: (2024)
Software-Defined Agentic Serving
di: Agarwal, Saurabh, et al.
Pubblicazione: (2026)
di: Agarwal, Saurabh, et al.
Pubblicazione: (2026)
DejaVu: A Minimalistic Mechanism for Distributed Plurality Consensus
di: d'Amore, Francesco, et al.
Pubblicazione: (2026)
di: d'Amore, Francesco, et al.
Pubblicazione: (2026)
Perpetual exploration in anonymous synchronous networks with a Byzantine black hole
di: Bhattacharya, Adri, et al.
Pubblicazione: (2025)
di: Bhattacharya, Adri, et al.
Pubblicazione: (2025)
Nalar: An agent serving framework
di: Laju, Marco, et al.
Pubblicazione: (2026)
di: Laju, Marco, et al.
Pubblicazione: (2026)
On the $h$-majority dynamics with many opinions
di: d'Amore, Francesco, et al.
Pubblicazione: (2025)
di: d'Amore, Francesco, et al.
Pubblicazione: (2025)
On a Voter Model with Context-Dependent Opinion Adoption
di: Becchetti, Luca, et al.
Pubblicazione: (2023)
di: Becchetti, Luca, et al.
Pubblicazione: (2023)
When Coordination Is Avoidable: A Monotonicity Analysis of Organizational Tasks
di: Ju, Harang
Pubblicazione: (2026)
di: Ju, Harang
Pubblicazione: (2026)
Cooperative Solutions to Exploration Tasks Under Speed and Budget Constraints
di: Karishma, et al.
Pubblicazione: (2022)
di: Karishma, et al.
Pubblicazione: (2022)
Asymptotic analysis of cooperative censoring policies in sensor networks
di: Fernandez-Bes, Jesus, et al.
Pubblicazione: (2025)
di: Fernandez-Bes, Jesus, et al.
Pubblicazione: (2025)
A Framework for Hybrid Collective Inference in Distributed Sensor Networks
di: Nash, Andrew, et al.
Pubblicazione: (2026)
di: Nash, Andrew, et al.
Pubblicazione: (2026)
Multitask Learning with Learned Task Relationships
di: Wan, Zirui, et al.
Pubblicazione: (2025)
di: Wan, Zirui, et al.
Pubblicazione: (2025)
Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches
di: Fang, Shaoke, et al.
Pubblicazione: (2026)
di: Fang, Shaoke, et al.
Pubblicazione: (2026)
Documenti analoghi
-
ScaleSim: Serving Large-Scale Multi-Agent Simulation with Invocation Distance-Based Memory Management
di: Pan, Zaifeng, et al.
Pubblicazione: (2026) -
Empowering Scientific Workflows with Federated Agents
di: Kamatar, Alok, et al.
Pubblicazione: (2025) -
FlashEvolve: Accelerating Agent Self-Evolution with Asynchronous Stage Orchestration
di: Hu, Zhengding, et al.
Pubblicazione: (2026) -
Aegis: Taxonomy and Optimizations for Overcoming Agent-Environment Failures in LLM Agents
di: Song, Kevin, et al.
Pubblicazione: (2025) -
Pythia: Exploiting Workflow Predictability for Efficient Agent-Native LLM Serving
di: Yu, Shan, et al.
Pubblicazione: (2026)