Semantic Scheduling for LLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Hua, Wenyue, Ding, Dujian, Gu, Yile, Ren, Yujie, Mei, Kai, Ma, Minghua, Wang, William Yang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models
by: Kamahori, Keisuke, et al.
Published: (2024)
by: Kamahori, Keisuke, et al.
Published: (2024)
Neuralink: Fast LLM Inference on Smartphones with Neuron Co-Activation Linking
by: Wang, Tuowei, et al.
Published: (2024)
by: Wang, Tuowei, et al.
Published: (2024)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
by: Feng, Shaoting, et al.
Published: (2025)
by: Feng, Shaoting, et al.
Published: (2025)
From Imperative to Declarative: Towards LLM-friendly OS Interfaces for Boosted Computer-Use Agents
by: Wang, Yuan, et al.
Published: (2025)
by: Wang, Yuan, et al.
Published: (2025)
ConsumerBench: Benchmarking Generative AI Applications on End-User Devices
by: Gu, Yile, et al.
Published: (2025)
by: Gu, Yile, et al.
Published: (2025)
AIOS: LLM Agent Operating System
by: Mei, Kai, et al.
Published: (2024)
by: Mei, Kai, et al.
Published: (2024)
Dynamic Speculative Agent Planning
by: Guan, Yilin, et al.
Published: (2025)
by: Guan, Yilin, et al.
Published: (2025)
AdaptCache: KV Cache Native Storage Hierarchy for Low-Delay and High-Quality Language Model Serving
by: Feng, Shaoting, et al.
Published: (2025)
by: Feng, Shaoting, et al.
Published: (2025)
Sawtooth Wavefront Reordering: Enhanced CuTile FlashAttention on NVIDIA GB10
by: Zhu, Yifan, et al.
Published: (2026)
by: Zhu, Yifan, et al.
Published: (2026)
OSWorld-Human: Benchmarking the Efficiency of Computer-Use Agents
by: Abhyankar, Reyna, et al.
Published: (2025)
by: Abhyankar, Reyna, et al.
Published: (2025)
An Integrated Artificial Intelligence Operating System for Advanced Low-Altitude Aviation Applications
by: Tan, Minzhe, et al.
Published: (2024)
by: Tan, Minzhe, et al.
Published: (2024)
Preparation Meets Opportunity: Enhancing Data Preprocessing for ML Training With Seneca
by: Desai, Omkar, et al.
Published: (2025)
by: Desai, Omkar, et al.
Published: (2025)
Towards Agentic OS: An LLM Agent Framework for Linux Schedulers
by: Zheng, Yusheng, et al.
Published: (2025)
by: Zheng, Yusheng, et al.
Published: (2025)
Enhancing Battery Storage Energy Arbitrage with Deep Reinforcement Learning and Time-Series Forecasting
by: Sage, Manuel, et al.
Published: (2024)
by: Sage, Manuel, et al.
Published: (2024)
FlexInfer: Breaking Memory Constraint via Flexible and Efficient Offloading for On-Device LLM Inference
by: Du, Hongchao, et al.
Published: (2025)
by: Du, Hongchao, et al.
Published: (2025)
Hardware-Assisted Virtualization of Neural Processing Units for Cloud Platforms
by: Xue, Yuqi, et al.
Published: (2024)
by: Xue, Yuqi, et al.
Published: (2024)
Samoyeds: Accelerating MoE Models with Structured Sparsity Leveraging Sparse Tensor Cores
by: Wu, Chenpeng, et al.
Published: (2025)
by: Wu, Chenpeng, et al.
Published: (2025)
FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation
by: Wu, Yinpeng, et al.
Published: (2026)
by: Wu, Yinpeng, et al.
Published: (2026)
Crab: A Semantics-Aware Checkpoint/Restore Runtime for Agent Sandboxes
by: Wu, Tianyuan, et al.
Published: (2026)
by: Wu, Tianyuan, et al.
Published: (2026)
Puzzle: Scheduling Multiple Deep Learning Models on Mobile Device with Heterogeneous Processors
by: Kang, Duseok, et al.
Published: (2025)
by: Kang, Duseok, et al.
Published: (2025)
Leveraging Machine Learning for Accurate IoT Device Identification in Dynamic Wireless Contexts
by: Tushir, Bhagyashri, et al.
Published: (2024)
by: Tushir, Bhagyashri, et al.
Published: (2024)
Cerebrum (AIOS SDK): A Platform for Agent Development, Deployment, Distribution, and Discovery
by: Rama, Balaji, et al.
Published: (2025)
by: Rama, Balaji, et al.
Published: (2025)
LatencyPrism: Online Non-intrusive Latency Sculpting for SLO-Guaranteed LLM Inference
by: Du, Yin, et al.
Published: (2026)
by: Du, Yin, et al.
Published: (2026)
Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference
by: Chu, Kexin, et al.
Published: (2025)
by: Chu, Kexin, et al.
Published: (2025)
NaSh: Guardrails for an LLM-Powered Natural Language Shell
by: Gyawali, Bimal Raj, et al.
Published: (2025)
by: Gyawali, Bimal Raj, et al.
Published: (2025)
Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
by: Li, Hanchen, et al.
Published: (2025)
by: Li, Hanchen, et al.
Published: (2025)
TempoNet: Slack-Quantized Transformer-Guided Reinforcement Scheduler for Adaptive Deadline-Centric Real-Time Dispatchs
by: Fu, Rong, et al.
Published: (2026)
by: Fu, Rong, et al.
Published: (2026)
SemaTune: Semantic-Aware Online OS Tuning with Large Language Models
by: Liargkovas, Georgios, et al.
Published: (2026)
by: Liargkovas, Georgios, et al.
Published: (2026)
SAGA: Workflow-Atomic Scheduling for AI Agent Inference on GPU Clusters
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
LiteCUA: Computer as MCP Server for Computer-Use Agent on AIOS
by: Mei, Kai, et al.
Published: (2025)
by: Mei, Kai, et al.
Published: (2025)
Diagnosing and Resolving Cloud Platform Instability with Multi-modal RAG LLMs
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
AgentCgroup: Understanding and Controlling OS Resources of AI Agents
by: Zheng, Yusheng, et al.
Published: (2026)
by: Zheng, Yusheng, et al.
Published: (2026)
ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented Generation
by: Wang, Shihao, et al.
Published: (2026)
by: Wang, Shihao, et al.
Published: (2026)
Token Management in Multi-Tenant AI Inference Platforms
by: Cunningham, William J.
Published: (2026)
by: Cunningham, William J.
Published: (2026)
The Missing Memory Hierarchy: Demand Paging for LLM Context Windows
by: Mason, Tony
Published: (2026)
by: Mason, Tony
Published: (2026)
TClone: Low-Latency Forking of Live GUI Environments for Computer-Use Agents
by: Huang, Yutong, et al.
Published: (2026)
by: Huang, Yutong, et al.
Published: (2026)
DeltaBox: Scaling Stateful AI Agents with Millisecond-Level Sandbox Checkpoint/Rollback
by: Dong, Yunpeng, et al.
Published: (2026)
by: Dong, Yunpeng, et al.
Published: (2026)
Composable OS Kernel Architectures for Autonomous Intelligence
by: Singh, Rajpreet, et al.
Published: (2025)
by: Singh, Rajpreet, et al.
Published: (2025)
Integrating Artificial Intelligence into Operating Systems: A Survey on Techniques, Applications, and Future Directions
by: Zhang, Yifan, et al.
Published: (2024)
by: Zhang, Yifan, et al.
Published: (2024)
Skim: Speculative Execution for Fast and Efficient Web Agents
by: Wong, Mike, et al.
Published: (2026)
by: Wong, Mike, et al.
Published: (2026)
Similar Items
-
Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models
by: Kamahori, Keisuke, et al.
Published: (2024) -
Neuralink: Fast LLM Inference on Smartphones with Neuron Co-Activation Linking
by: Wang, Tuowei, et al.
Published: (2024) -
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
by: Feng, Shaoting, et al.
Published: (2025) -
From Imperative to Declarative: Towards LLM-friendly OS Interfaces for Boosted Computer-Use Agents
by: Wang, Yuan, et al.
Published: (2025) -
ConsumerBench: Benchmarking Generative AI Applications on End-User Devices
by: Gu, Yile, et al.
Published: (2025)