AMV-L: Lifecycle-Managed Agent Memory for Tail-Latency Control in Long-Running LLM Systems
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Bamidele, Emmanuel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents
von: Wang, Taiyi, et al.
Veröffentlicht: (2024)
von: Wang, Taiyi, et al.
Veröffentlicht: (2024)
Soar: Design and Deployment of A Smart Roadside Infrastructure System for Autonomous Driving
von: Shi, Shuyao, et al.
Veröffentlicht: (2024)
von: Shi, Shuyao, et al.
Veröffentlicht: (2024)
An Uncertainty-Aware Resilience Micro-Agent for Causal Observability in the Computing Continuum
von: De Silva, Suvi, et al.
Veröffentlicht: (2026)
von: De Silva, Suvi, et al.
Veröffentlicht: (2026)
Improving accuracy and convergence of federated learning edge computing methods for generalized DER forecasting applications in power grid
von: Nair, Vineet Jagadeesan, et al.
Veröffentlicht: (2024)
von: Nair, Vineet Jagadeesan, et al.
Veröffentlicht: (2024)
Evaluation of a Foundational Model and Stochastic Models for Forecasting Sporadic or Spiky Production Outages of High-Performance Machine Learning Services
von: Yim, Keun Soo
Veröffentlicht: (2025)
von: Yim, Keun Soo
Veröffentlicht: (2025)
VREM-FL: Mobility-Aware Computation-Scheduling Co-Design for Vehicular Federated Learning
von: Ballotta, Luca, et al.
Veröffentlicht: (2023)
von: Ballotta, Luca, et al.
Veröffentlicht: (2023)
Fully Decentralized Joint Learning of Personalized Models and Collaboration Graphs
von: Zantedeschi, Valentina, et al.
Veröffentlicht: (2019)
von: Zantedeschi, Valentina, et al.
Veröffentlicht: (2019)
Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints
von: Ao, Ruicheng, et al.
Veröffentlicht: (2025)
von: Ao, Ruicheng, et al.
Veröffentlicht: (2025)
Taming Latency-Memory Trade-Off in MoE-Based LLM Serving via Fine-Grained Expert Offloading
von: Yu, Hanfei, et al.
Veröffentlicht: (2025)
von: Yu, Hanfei, et al.
Veröffentlicht: (2025)
Lifelong Federated Reinforcement Learning: A Learning Architecture for Navigation in Cloud Robotic Systems
von: Liu, Boyi, et al.
Veröffentlicht: (2019)
von: Liu, Boyi, et al.
Veröffentlicht: (2019)
The Smart Buildings Control Suite: A Diverse Open Source Benchmark to Evaluate and Scale HVAC Control Policies for Sustainability
von: Goldfeder, Judah, et al.
Veröffentlicht: (2024)
von: Goldfeder, Judah, et al.
Veröffentlicht: (2024)
Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter
von: Hu, Qinghao, et al.
Veröffentlicht: (2025)
von: Hu, Qinghao, et al.
Veröffentlicht: (2025)
SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips
von: Yu, Jiahuan, et al.
Veröffentlicht: (2026)
von: Yu, Jiahuan, et al.
Veröffentlicht: (2026)
Input-Based Ensemble-Learning Method for Dynamic Memory Configuration of Serverless Computing Functions
von: Agarwal, Siddharth, et al.
Veröffentlicht: (2024)
von: Agarwal, Siddharth, et al.
Veröffentlicht: (2024)
Reliable Microservice Tail Latency Prediction via Decoupled Dual-Stream Learning and Gradient Modulation
von: Qian, Wenzhuo, et al.
Veröffentlicht: (2025)
von: Qian, Wenzhuo, et al.
Veröffentlicht: (2025)
Win Fast or Lose Slow: Balancing Speed and Accuracy in Latency-Sensitive Decisions of LLMs
von: Kang, Hao, et al.
Veröffentlicht: (2025)
von: Kang, Hao, et al.
Veröffentlicht: (2025)
Adaptive Workload Distribution for Accuracy-aware DNN Inference on Collaborative Edge Platforms
von: Taufique, Zain, et al.
Veröffentlicht: (2023)
von: Taufique, Zain, et al.
Veröffentlicht: (2023)
Equilibrium in the Computing Continuum through Active Inference
von: Sedlak, Boris, et al.
Veröffentlicht: (2023)
von: Sedlak, Boris, et al.
Veröffentlicht: (2023)
Federated reinforcement learning for robot motion planning with zero-shot generalization
von: Yuan, Zhenyuan, et al.
Veröffentlicht: (2024)
von: Yuan, Zhenyuan, et al.
Veröffentlicht: (2024)
Sustainability of Data Center Digital Twins with Reinforcement Learning
von: Sarkar, Soumyendu, et al.
Veröffentlicht: (2024)
von: Sarkar, Soumyendu, et al.
Veröffentlicht: (2024)
Accuracy-Delay Trade-Off in LLM Offloading via Token-Level Uncertainty
von: Kim, Yumin, et al.
Veröffentlicht: (2026)
von: Kim, Yumin, et al.
Veröffentlicht: (2026)
S-Bus: Automatic Read-Set Reconstruction for Multi-Agent LLM State Coordination
von: Khan, Sajjad
Veröffentlicht: (2026)
von: Khan, Sajjad
Veröffentlicht: (2026)
Prompt-Aware Scheduling for Low-Latency LLM Serving
von: Tao, Yiheng, et al.
Veröffentlicht: (2025)
von: Tao, Yiheng, et al.
Veröffentlicht: (2025)
JITServe: SLO-aware LLM Serving with Imprecise Request Information
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
Systemic approach for modeling a generic smart grid
von: Amor, Sofiane Ben, et al.
Veröffentlicht: (2025)
von: Amor, Sofiane Ben, et al.
Veröffentlicht: (2025)
SpecKV: Adaptive Speculative Decoding with Compression-Aware Gamma Selection
von: Shukla, Shikhar
Veröffentlicht: (2026)
von: Shukla, Shikhar
Veröffentlicht: (2026)
Technical Report on Reinforcement Learning Control on the Lucas-Nülle Inverted Pendulum
von: Schenke, Maximilian, et al.
Veröffentlicht: (2024)
von: Schenke, Maximilian, et al.
Veröffentlicht: (2024)
DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training
von: Li, Dacheng, et al.
Veröffentlicht: (2023)
von: Li, Dacheng, et al.
Veröffentlicht: (2023)
CommFuse: Hiding Tail Latency via Communication Decomposition and Fusion for Distributed LLM Training
von: Karim, Rezaul, et al.
Veröffentlicht: (2026)
von: Karim, Rezaul, et al.
Veröffentlicht: (2026)
LatencyPrism: Online Non-intrusive Latency Sculpting for SLO-Guaranteed LLM Inference
von: Du, Yin, et al.
Veröffentlicht: (2026)
von: Du, Yin, et al.
Veröffentlicht: (2026)
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
von: Jiang, Xuanlin, et al.
Veröffentlicht: (2024)
von: Jiang, Xuanlin, et al.
Veröffentlicht: (2024)
Rudder: Steering Prefetching in Distributed GNN Training using LLM Agents
von: Sarkar, Aishwarya, et al.
Veröffentlicht: (2026)
von: Sarkar, Aishwarya, et al.
Veröffentlicht: (2026)
Secure Cluster-Based Hierarchical Federated Learning in Vehicular Networks
von: HaghighiFard, M. Saeid, et al.
Veröffentlicht: (2025)
von: HaghighiFard, M. Saeid, et al.
Veröffentlicht: (2025)
Non-Convex Over-the-Air Heterogeneous Federated Learning: A Bias-Variance Trade-off
von: Abrar, Muhammad Faraz Ul, et al.
Veröffentlicht: (2025)
von: Abrar, Muhammad Faraz Ul, et al.
Veröffentlicht: (2025)
Hierarchical Federated ADMM
von: Azimi-Abarghouyi, Seyed Mohammad, et al.
Veröffentlicht: (2024)
von: Azimi-Abarghouyi, Seyed Mohammad, et al.
Veröffentlicht: (2024)
Autellix: An Efficient Serving Engine for LLM Agents as General Programs
von: Luo, Michael, et al.
Veröffentlicht: (2025)
von: Luo, Michael, et al.
Veröffentlicht: (2025)
Llamas on the Web: Memory-Efficient, Performance-Portable, and Multi-Precision LLM Inference with WebGPU
von: Levine, Reese, et al.
Veröffentlicht: (2026)
von: Levine, Reese, et al.
Veröffentlicht: (2026)
Ensuring Fair LLM Serving Amid Diverse Applications
von: Khan, Redwan Ibne Seraj, et al.
Veröffentlicht: (2024)
von: Khan, Redwan Ibne Seraj, et al.
Veröffentlicht: (2024)
Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference
von: Deshmukh, Dhruv, et al.
Veröffentlicht: (2025)
von: Deshmukh, Dhruv, et al.
Veröffentlicht: (2025)
FL-GUARD: A Holistic Framework for Run-Time Detection and Recovery of Negative Federated Learning
von: Lin, Hong, et al.
Veröffentlicht: (2024)
von: Lin, Hong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents
von: Wang, Taiyi, et al.
Veröffentlicht: (2024) -
Soar: Design and Deployment of A Smart Roadside Infrastructure System for Autonomous Driving
von: Shi, Shuyao, et al.
Veröffentlicht: (2024) -
An Uncertainty-Aware Resilience Micro-Agent for Causal Observability in the Computing Continuum
von: De Silva, Suvi, et al.
Veröffentlicht: (2026) -
Improving accuracy and convergence of federated learning edge computing methods for generalized DER forecasting applications in power grid
von: Nair, Vineet Jagadeesan, et al.
Veröffentlicht: (2024) -
Evaluation of a Foundational Model and Stochastic Models for Forecasting Sporadic or Spiky Production Outages of High-Performance Machine Learning Services
von: Yim, Keun Soo
Veröffentlicht: (2025)