Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Hanchen, He, Runyuan, Mang, Qiuyang, Zhang, Qizheng, Mao, Huanzhi, Chen, Xiaokun, Zhou, Hangrui, Cheung, Alvin, Gonzalez, Joseph, Stoica, Ion |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Revisiting Cache Freshness for Emerging Real-Time Applications
por: Mao, Ziming, et al.
Publicado: (2024)
por: Mao, Ziming, et al.
Publicado: (2024)
Towards More Economical Context-Augmented LLM Generation by Reusing Stored KV Cache
por: Li, Hanchen, et al.
Publicado: (2025)
por: Li, Hanchen, et al.
Publicado: (2025)
An Online Gradient-Based Caching Policy with Logarithmic Complexity and Regret Guarantees
por: Carra, Damiano, et al.
Publicado: (2024)
por: Carra, Damiano, et al.
Publicado: (2024)
CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving
por: Liu, Yuhan, et al.
Publicado: (2023)
por: Liu, Yuhan, et al.
Publicado: (2023)
The NIC should be part of the OS
por: Xu, Pengcheng, et al.
Publicado: (2025)
por: Xu, Pengcheng, et al.
Publicado: (2025)
Supporting Deterministic Traffic on Standard NICs
por: Xue, Chuanyu, et al.
Publicado: (2025)
por: Xue, Chuanyu, et al.
Publicado: (2025)
Towards Timing Isolation for Mixed-Criticality Communication in Software-Defined Vehicles
por: Meszlényi, Lóránt, et al.
Publicado: (2025)
por: Meszlényi, Lóránt, et al.
Publicado: (2025)
Characterizing Network Requirements for GPU API Remoting in AI Applications
por: Wang, Tianxia, et al.
Publicado: (2024)
por: Wang, Tianxia, et al.
Publicado: (2024)
Chamelio: A Fast Shared Cloud Network Stack for Isolated Tenant-Defined Protocols
por: Stolet, Matheus, et al.
Publicado: (2026)
por: Stolet, Matheus, et al.
Publicado: (2026)
Tail Contagion: Sub-microsecond Time Protection in Shared Software Network Datapaths
por: Stolet, Matheus, et al.
Publicado: (2023)
por: Stolet, Matheus, et al.
Publicado: (2023)
On Measuring Available Capacity in High-speed Cloud Networks
por: Madanagopal, Ganapathy Raman, et al.
Publicado: (2025)
por: Madanagopal, Ganapathy Raman, et al.
Publicado: (2025)
Exploring Busy Period for Worst-Case Deadline Failure Probability Analysis
por: Liu, Junyi, et al.
Publicado: (2025)
por: Liu, Junyi, et al.
Publicado: (2025)
Söze: One Network Telemetry Is All You Need for Per-flow Weighted Bandwidth Allocation at Scale
por: Wang, Weitao, et al.
Publicado: (2025)
por: Wang, Weitao, et al.
Publicado: (2025)
Scaling Data Center TCP to Terabits with Laminar
por: Shashidhara, Rajath, et al.
Publicado: (2025)
por: Shashidhara, Rajath, et al.
Publicado: (2025)
Saving Storage Space Using Files on the Web
por: Saric, Kevin, et al.
Publicado: (2025)
por: Saric, Kevin, et al.
Publicado: (2025)
FlexBSO: Flexible Block Storage Offload for Datacenters
por: Aschenbrenner, Vojtech, et al.
Publicado: (2024)
por: Aschenbrenner, Vojtech, et al.
Publicado: (2024)
EDM: An Ultra-Low Latency Ethernet Fabric for Memory Disaggregation
por: Su, Weigao, et al.
Publicado: (2024)
por: Su, Weigao, et al.
Publicado: (2024)
uTNT: Unikernels for Efficient and Flexible Internet Probing
por: Letemple, Maxime, et al.
Publicado: (2024)
por: Letemple, Maxime, et al.
Publicado: (2024)
ONCache: A Cache-Based Low-Overhead Container Overlay Network
por: Lin, Shengkai, et al.
Publicado: (2023)
por: Lin, Shengkai, et al.
Publicado: (2023)
Understanding and Enhancing Linux Kernel-based Packet Switching on WiFi Access Points
por: Zhang, Shiqi, et al.
Publicado: (2024)
por: Zhang, Shiqi, et al.
Publicado: (2024)
Sensifi: A Wireless Sensing System for Ultra-High-Rate Applications
por: Li, Chia-Chi, et al.
Publicado: (2020)
por: Li, Chia-Chi, et al.
Publicado: (2020)
Energy-Aware CPU Orchestration in O-RAN: A dApp-Driven Lightweight Approach
por: Crespo, Francisco, et al.
Publicado: (2025)
por: Crespo, Francisco, et al.
Publicado: (2025)
SPECTRE: A Hybrid System for an Adaptative and Optimised Cyber Threats Detection, Response and Investigation in Volatile Memory
por: Syed, Arslan Tariq, et al.
Publicado: (2025)
por: Syed, Arslan Tariq, et al.
Publicado: (2025)
Investigation of Advanced Persistent Threats Network-based Tactics, Techniques and Procedures
por: Alageel, Almuthanna, et al.
Publicado: (2025)
por: Alageel, Almuthanna, et al.
Publicado: (2025)
GPUs, CPUs, and... NICs: Rethinking the Network's Role in Serving Complex AI Pipelines
por: Wong, Mike, et al.
Publicado: (2025)
por: Wong, Mike, et al.
Publicado: (2025)
Qoala: an Application Execution Environment for Quantum Internet Nodes
por: van der Vecht, Bart, et al.
Publicado: (2025)
por: van der Vecht, Bart, et al.
Publicado: (2025)
Fast Networks for High-Performance Distributed Trust
por: Liu, Yicheng, et al.
Publicado: (2025)
por: Liu, Yicheng, et al.
Publicado: (2025)
Design and demonstration of an operating system for executing applications on quantum network nodes
por: Donne, Carlo Delle, et al.
Publicado: (2024)
por: Donne, Carlo Delle, et al.
Publicado: (2024)
HookChain: A new perspective for Bypassing EDR Solutions
por: Junior, Helvio Carvalho
Publicado: (2024)
por: Junior, Helvio Carvalho
Publicado: (2024)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
por: Feng, Shaoting, et al.
Publicado: (2025)
por: Feng, Shaoting, et al.
Publicado: (2025)
SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference
por: Liu, Hongyao, et al.
Publicado: (2026)
por: Liu, Hongyao, et al.
Publicado: (2026)
bypass4netns: Accelerating TCP/IP Communications in Rootless Containers
por: Matsumoto, Naoki, et al.
Publicado: (2024)
por: Matsumoto, Naoki, et al.
Publicado: (2024)
Deterministic and Probabilistic P4-Enabled Lightweight In-Band Network Telemetry
por: Papadopoulos, Konstantinos, et al.
Publicado: (2024)
por: Papadopoulos, Konstantinos, et al.
Publicado: (2024)
Ultra Ethernet's Design Principles and Architectural Innovations
por: Hoefler, Torsten, et al.
Publicado: (2025)
por: Hoefler, Torsten, et al.
Publicado: (2025)
Accelerator-as-a-Service in Public Clouds: An Intra-Host Traffic Management View for Performance Isolation in the Wild
por: Zhao, Jiechen, et al.
Publicado: (2024)
por: Zhao, Jiechen, et al.
Publicado: (2024)
Leveraging Machine Learning for Accurate IoT Device Identification in Dynamic Wireless Contexts
por: Tushir, Bhagyashri, et al.
Publicado: (2024)
por: Tushir, Bhagyashri, et al.
Publicado: (2024)
Boxer: FaaSt Ephemeral Elasticity for Off-the-Shelf Cloud Applications
por: Wawrzoniak, Michael, et al.
Publicado: (2024)
por: Wawrzoniak, Michael, et al.
Publicado: (2024)
Fast Userspace Networking for the Rest of Us
por: Sanaee, Alireza, et al.
Publicado: (2025)
por: Sanaee, Alireza, et al.
Publicado: (2025)
Imaginary Machines: A Serverless Model for Cloud Applications
por: Wawrzoniak, Michael, et al.
Publicado: (2024)
por: Wawrzoniak, Michael, et al.
Publicado: (2024)
Hazel: Secure and Efficient Disaggregated Storage
por: Chrapek, Marcin, et al.
Publicado: (2025)
por: Chrapek, Marcin, et al.
Publicado: (2025)
Ejemplares similares
-
Revisiting Cache Freshness for Emerging Real-Time Applications
por: Mao, Ziming, et al.
Publicado: (2024) -
Towards More Economical Context-Augmented LLM Generation by Reusing Stored KV Cache
por: Li, Hanchen, et al.
Publicado: (2025) -
An Online Gradient-Based Caching Policy with Logarithmic Complexity and Regret Guarantees
por: Carra, Damiano, et al.
Publicado: (2024) -
CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving
por: Liu, Yuhan, et al.
Publicado: (2023) -
The NIC should be part of the OS
por: Xu, Pengcheng, et al.
Publicado: (2025)