Idleness is Relative: Exploiting Tool-Call Idle Windows for Offloading in Agentic Systems with MORI
Fuente:
arXiv
Saved in:
| Main Authors: | Xia, Tian, Li, Hanchen, Li, Zhifei, Chen, Xiaokun, Kang, Hao, Qiao, Yifan, Xu, Yi, Stoica, Ion |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
by: Li, Hanchen, et al.
Published: (2025)
by: Li, Hanchen, et al.
Published: (2025)
Towards Efficient and Practical GPU Multitasking in the Era of LLM
by: Xing, Jiarong, et al.
Published: (2025)
by: Xing, Jiarong, et al.
Published: (2025)
HeteroPod: XPU-Accelerated Infrastructure Offloading for Commodity Cloud-Native Applications
by: Yang, Bicheng, et al.
Published: (2025)
by: Yang, Bicheng, et al.
Published: (2025)
ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System
by: Kang, Hao, et al.
Published: (2026)
by: Kang, Hao, et al.
Published: (2026)
Revisiting Cache Freshness for Emerging Real-Time Applications
by: Mao, Ziming, et al.
Published: (2024)
by: Mao, Ziming, et al.
Published: (2024)
Dissecting CXL Memory Performance at Scale: Analysis, Modeling, and Optimization
by: Liu, Jinshu, et al.
Published: (2024)
by: Liu, Jinshu, et al.
Published: (2024)
Rethinking Inter-Process Communication with Memory Operation Offloading
by: Park, Misun, et al.
Published: (2026)
by: Park, Misun, et al.
Published: (2026)
Wave: Offloading Resource Management to SmartNIC Cores
by: Humphries, Jack Tigar, et al.
Published: (2024)
by: Humphries, Jack Tigar, et al.
Published: (2024)
Exploiting Dependency and Parallelism: Real-Time Scheduling and Analysis for GPU Tasks
by: Zhang, Yuanhai, et al.
Published: (2026)
by: Zhang, Yuanhai, et al.
Published: (2026)
HALO: A Fine-Grained Resource Sharing Quantum Operating System
by: Ye, John Zhuoyang, et al.
Published: (2026)
by: Ye, John Zhuoyang, et al.
Published: (2026)
B-Side: Binary-Level Static System Call Identification
by: Thévenon, Gaspard, et al.
Published: (2024)
by: Thévenon, Gaspard, et al.
Published: (2024)
Exploiting Application-to-Architecture Dependencies for Designing Scalable OS
by: Xiao, Yao, et al.
Published: (2025)
by: Xiao, Yao, et al.
Published: (2025)
Inspection of I/O Operations from System Call Traces using Directly-Follows-Graph
by: Sankaran, Aravind, et al.
Published: (2024)
by: Sankaran, Aravind, et al.
Published: (2024)
Foreactor: Exploiting Storage I/O Parallelism with Explicit Speculation
by: Hu, Guanzhou, et al.
Published: (2024)
by: Hu, Guanzhou, et al.
Published: (2024)
RAGDoll: Efficient Offloading-based Online RAG System on a Single GPU
by: Yu, Weiping, et al.
Published: (2025)
by: Yu, Weiping, et al.
Published: (2025)
Tracial states on groupoid $C^*$-algebras and essential freeness
by: Li, Kang, et al.
Published: (2024)
by: Li, Kang, et al.
Published: (2024)
Coach: Exploiting Temporal Patterns for All-Resource Oversubscription in Cloud Platforms
by: Reidys, Benjamin, et al.
Published: (2025)
by: Reidys, Benjamin, et al.
Published: (2025)
ByteFS: System Support for (CXL-based) Memory-Semantic Solid-State Drives
by: Li, Shaobo, et al.
Published: (2025)
by: Li, Shaobo, et al.
Published: (2025)
FlexBSO: Flexible Block Storage Offload for Datacenters
by: Aschenbrenner, Vojtech, et al.
Published: (2024)
by: Aschenbrenner, Vojtech, et al.
Published: (2024)
numaPTE: Managing Page-Tables and TLBs on NUMA Systems
by: Gao, Bin, et al.
Published: (2024)
by: Gao, Bin, et al.
Published: (2024)
Exploiting Page Faults for Covert Communication
by: Swaminathan, Sathvik
Published: (2025)
by: Swaminathan, Sathvik
Published: (2025)
Toward Systems Foundations for Agentic Exploration
by: Xu, Jiakai, et al.
Published: (2025)
by: Xu, Jiakai, et al.
Published: (2025)
Tensor free probability theory: asymptotic tensor freeness and central limit theorem
by: Nechita, Ion, et al.
Published: (2025)
by: Nechita, Ion, et al.
Published: (2025)
FlexInfer: Breaking Memory Constraint via Flexible and Efficient Offloading for On-Device LLM Inference
by: Du, Hongchao, et al.
Published: (2025)
by: Du, Hongchao, et al.
Published: (2025)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
by: Feng, Shaoting, et al.
Published: (2025)
by: Feng, Shaoting, et al.
Published: (2025)
CHRONOS: Compensating Hardware Related Overheads with Native Multi Timer Support for Real-Time Operating Systems
by: Heider, Kay, et al.
Published: (2025)
by: Heider, Kay, et al.
Published: (2025)
Integrating Artificial Intelligence into Operating Systems: A Survey on Techniques, Applications, and Future Directions
by: Zhang, Yifan, et al.
Published: (2024)
by: Zhang, Yifan, et al.
Published: (2024)
DFUSE: Strongly Consistent Write-Back Kernel Caching for Distributed Userspace File Systems
by: Li, Haoyu, et al.
Published: (2025)
by: Li, Haoyu, et al.
Published: (2025)
XBOF: A Cost-Efficient CXL JBOF with Inter-SSD Compute Resource Sharing
by: Yi, Shushu, et al.
Published: (2025)
by: Yi, Shushu, et al.
Published: (2025)
Non-free almost finite actions for locally finite-by-virtually $\mathbb{Z}$ groups
by: Li, Kang, et al.
Published: (2023)
by: Li, Kang, et al.
Published: (2023)
Nexus: Transparent I/O Offloading for High-Density Serverless Computing
by: Park, JooYoung, et al.
Published: (2026)
by: Park, JooYoung, et al.
Published: (2026)
Nova: Real-Time Agentic Vision-Language Model Serving with Adaptive Cross-Stage Parallelization
by: Xu, Yuhang, et al.
Published: (2025)
by: Xu, Yuhang, et al.
Published: (2025)
LLM as a System Service on Mobile Devices
by: Yin, Wangsong, et al.
Published: (2024)
by: Yin, Wangsong, et al.
Published: (2024)
Columbo: Low Level End-to-End System Traces through Modular Full-System Simulation
by: Görgen, Jakob, et al.
Published: (2024)
by: Görgen, Jakob, et al.
Published: (2024)
Information sharing and procurement strategies in agricultural supply chains under retailer competition
by: Yaping Zhao, et al.
Published: (2025)
by: Yaping Zhao, et al.
Published: (2025)
Hybrid Adaptive Tuning for Tiered Memory Systems
by: Wang, Xi, et al.
Published: (2026)
by: Wang, Xi, et al.
Published: (2026)
Way-below Relation and Tensor Products
by: Ivanescu, Cristian, et al.
Published: (2025)
by: Ivanescu, Cristian, et al.
Published: (2025)
Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving
by: Qiu, Shi, et al.
Published: (2026)
by: Qiu, Shi, et al.
Published: (2026)
SARA: A Stall-Aware Memory Allocation Strategy for Mixed-Criticality Systems
by: Lee, Meng-Chia, et al.
Published: (2025)
by: Lee, Meng-Chia, et al.
Published: (2025)
PipeANN-Filter: An Efficient Filtered Vector Search System on SSD
by: Guo, Hao, et al.
Published: (2026)
by: Guo, Hao, et al.
Published: (2026)
Similar Items
-
Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
by: Li, Hanchen, et al.
Published: (2025) -
Towards Efficient and Practical GPU Multitasking in the Era of LLM
by: Xing, Jiarong, et al.
Published: (2025) -
HeteroPod: XPU-Accelerated Infrastructure Offloading for Commodity Cloud-Native Applications
by: Yang, Bicheng, et al.
Published: (2025) -
ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System
by: Kang, Hao, et al.
Published: (2026) -
Revisiting Cache Freshness for Emerging Real-Time Applications
by: Mao, Ziming, et al.
Published: (2024)