Saved in:
| Main Authors: | Wang, Zhibin, Zhong, Ziyu, Shen, Nuo, Zhou, Yuhang, Gu, Rong, Zhong, Sheng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.19893 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Combining Type Checking and Formal Verification for Lightweight OS Correctness
by: Ijaz, Ramla, et al.
Published: (2024)
by: Ijaz, Ramla, et al.
Published: (2024)
RTP-LLM: High-Performance Alibaba LLM Inference Engine
by: Tan, Boyu, et al.
Published: (2026)
by: Tan, Boyu, et al.
Published: (2026)
Skim: Speculative Execution for Fast and Efficient Web Agents
by: Wong, Mike, et al.
Published: (2026)
by: Wong, Mike, et al.
Published: (2026)
PhoenixOS: Concurrent OS-level GPU Checkpoint and Restore with Validated Speculation
by: Wei, Xingda, et al.
Published: (2024)
by: Wei, Xingda, et al.
Published: (2024)
Semantic Scheduling for LLM Inference
by: Hua, Wenyue, et al.
Published: (2025)
by: Hua, Wenyue, et al.
Published: (2025)
Foreactor: Exploiting Storage I/O Parallelism with Explicit Speculation
by: Hu, Guanzhou, et al.
Published: (2024)
by: Hu, Guanzhou, et al.
Published: (2024)
Talyxion: From Speculation to Optimization in Risk Managed Crypto Portfolio Allocation
by: Nguyen, Thanh
Published: (2025)
by: Nguyen, Thanh
Published: (2025)
FlexInfer: Breaking Memory Constraint via Flexible and Efficient Offloading for On-Device LLM Inference
by: Du, Hongchao, et al.
Published: (2025)
by: Du, Hongchao, et al.
Published: (2025)
Towards Fully-fledged GPU Multitasking via Proactive Memory Scheduling
by: Shen, Weihang, et al.
Published: (2025)
by: Shen, Weihang, et al.
Published: (2025)
Holistic Heterogeneous Scheduling for Autonomous Applications using Fine-grained, Multi-XPU Abstraction
by: Han, Mingcong, et al.
Published: (2025)
by: Han, Mingcong, et al.
Published: (2025)
C2CServe: Leveraging NVLink-C2C for Elastic Serverless LLM Serving on MIG
by: Luo, Shutian, et al.
Published: (2026)
by: Luo, Shutian, et al.
Published: (2026)
GoCkpt: Gradient-Assisted Multi-Step overlapped Checkpointing for Efficient LLM Training
by: Zhang, Keyao, et al.
Published: (2025)
by: Zhang, Keyao, et al.
Published: (2025)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
by: Feng, Shaoting, et al.
Published: (2025)
by: Feng, Shaoting, et al.
Published: (2025)
XBOF: A Cost-Efficient CXL JBOF with Inter-SSD Compute Resource Sharing
by: Yi, Shushu, et al.
Published: (2025)
by: Yi, Shushu, et al.
Published: (2025)
LMetric: Simple is Better - Multiplication May Be All You Need for LLM Request Scheduling
by: Zhang, Dingyan, et al.
Published: (2026)
by: Zhang, Dingyan, et al.
Published: (2026)
Dissecting the Impact of Mobile DVFS Governors on LLM Inference Performance and Energy Efficiency
by: Zhang, Zongpu, et al.
Published: (2025)
by: Zhang, Zongpu, et al.
Published: (2025)
Blindfold: Confidential Memory Management by Untrusted Operating System
by: Li, Caihua, et al.
Published: (2024)
by: Li, Caihua, et al.
Published: (2024)
Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference
by: Chu, Kexin, et al.
Published: (2025)
by: Chu, Kexin, et al.
Published: (2025)
My CXL Pool Obviates Your PCIe Switch
by: Zhong, Yuhong, et al.
Published: (2025)
by: Zhong, Yuhong, et al.
Published: (2025)
Nova: Real-Time Agentic Vision-Language Model Serving with Adaptive Cross-Stage Parallelization
by: Xu, Yuhang, et al.
Published: (2025)
by: Xu, Yuhang, et al.
Published: (2025)
Don't Let AI Agents YOLO Your Files: Shifting Information and Control to Filesystems for Agent Safety and Autonomy
by: Zhong, Shawn Wanxiang, et al.
Published: (2026)
by: Zhong, Shawn Wanxiang, et al.
Published: (2026)
Towards High-Goodput LLM Serving with Prefill-decode Multiplexing
by: Chen, Yukang, et al.
Published: (2025)
by: Chen, Yukang, et al.
Published: (2025)
ParaCell: Paravirtualized Secure Containers with Lightweight Intra-Container Isolation and Intent-Driven Memory Management
by: Wu, Yiyang, et al.
Published: (2026)
by: Wu, Yiyang, et al.
Published: (2026)
ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System
by: Kang, Hao, et al.
Published: (2026)
by: Kang, Hao, et al.
Published: (2026)
Ringmaster: How to juggle high-throughput host OS system calls from TrustZone TEEs
by: Habeeb, Richard, et al.
Published: (2026)
by: Habeeb, Richard, et al.
Published: (2026)
Characterizing Network Requirements for GPU API Remoting in AI Applications
by: Wang, Tianxia, et al.
Published: (2024)
by: Wang, Tianxia, et al.
Published: (2024)
Towards Efficient and Practical GPU Multitasking in the Era of LLM
by: Xing, Jiarong, et al.
Published: (2025)
by: Xing, Jiarong, et al.
Published: (2025)
Valve: Production Online-Offline Inference Colocation with Jointly-Bounded Preemption Latency and Rate
by: Liu, Fangyue, et al.
Published: (2026)
by: Liu, Fangyue, et al.
Published: (2026)
Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving
by: Qiu, Shi, et al.
Published: (2026)
by: Qiu, Shi, et al.
Published: (2026)
LLM as a System Service on Mobile Devices
by: Yin, Wangsong, et al.
Published: (2024)
by: Yin, Wangsong, et al.
Published: (2024)
LatencyPrism: Online Non-intrusive Latency Sculpting for SLO-Guaranteed LLM Inference
by: Du, Yin, et al.
Published: (2026)
by: Du, Yin, et al.
Published: (2026)
RUISA Operational Ecosystem Architecture
by: AL Mohtar, Mouayad
Published: (2026)
by: AL Mohtar, Mouayad
Published: (2026)
Energy-Efficient Computation with DVFS using Deep Reinforcement Learning for Multi-Task Systems in Edge Computing
by: Li, Xinyi, et al.
Published: (2024)
by: Li, Xinyi, et al.
Published: (2024)
Testing Access-Control Configuration Changes for Web Applications
by: Xiang, Chengcheng, et al.
Published: (2025)
by: Xiang, Chengcheng, et al.
Published: (2025)
Neuralink: Fast LLM Inference on Smartphones with Neuron Co-Activation Linking
by: Wang, Tuowei, et al.
Published: (2024)
by: Wang, Tuowei, et al.
Published: (2024)
AgentRM: An OS-Inspired Resource Manager for LLM Agent Systems
by: She, Jianshu
Published: (2026)
by: She, Jianshu
Published: (2026)
Efficient Memory Tiering in a Virtual Machine
by: Prakash, Chandra, et al.
Published: (2025)
by: Prakash, Chandra, et al.
Published: (2025)
FusionANNS: An Efficient CPU/GPU Cooperative Processing Architecture for Billion-scale Approximate Nearest Neighbor Search
by: Tian, Bing, et al.
Published: (2024)
by: Tian, Bing, et al.
Published: (2024)
Data-driven Software-based Power Estimation for Embedded Devices
by: Wang, Haoyu, et al.
Published: (2024)
by: Wang, Haoyu, et al.
Published: (2024)
Flare: Anomaly Diagnostics for Divergent LLM Training in GPU Clusters of Thousand-Plus Scale
by: Cui, Weihao, et al.
Published: (2025)
by: Cui, Weihao, et al.
Published: (2025)
Similar Items
-
Combining Type Checking and Formal Verification for Lightweight OS Correctness
by: Ijaz, Ramla, et al.
Published: (2024) -
RTP-LLM: High-Performance Alibaba LLM Inference Engine
by: Tan, Boyu, et al.
Published: (2026) -
Skim: Speculative Execution for Fast and Efficient Web Agents
by: Wong, Mike, et al.
Published: (2026) -
PhoenixOS: Concurrent OS-level GPU Checkpoint and Restore with Validated Speculation
by: Wei, Xingda, et al.
Published: (2024) -
Semantic Scheduling for LLM Inference
by: Hua, Wenyue, et al.
Published: (2025)