Towards Fully-fledged GPU Multitasking via Proactive Memory Scheduling
Fuente:
arXiv
Saved in:
| Main Authors: | Shen, Weihang, Chen, Yinqiu, Chen, Rong, Chen, Haibo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Holistic Heterogeneous Scheduling for Autonomous Applications using Fine-grained, Multi-XPU Abstraction
by: Han, Mingcong, et al.
Published: (2025)
by: Han, Mingcong, et al.
Published: (2025)
Microsecond-scale Dynamic Validation of Idempotency for GPU Kernels
by: Han, Mingcong, et al.
Published: (2024)
by: Han, Mingcong, et al.
Published: (2024)
Characterizing Network Requirements for GPU API Remoting in AI Applications
by: Wang, Tianxia, et al.
Published: (2024)
by: Wang, Tianxia, et al.
Published: (2024)
Towards Efficient and Practical GPU Multitasking in the Era of LLM
by: Xing, Jiarong, et al.
Published: (2025)
by: Xing, Jiarong, et al.
Published: (2025)
Optimizing Tree-structure Indexes for CXL-based Heterogeneous Memory with SINLK
by: Zhao, Haoru, et al.
Published: (2025)
by: Zhao, Haoru, et al.
Published: (2025)
ParaCell: Paravirtualized Secure Containers with Lightweight Intra-Container Isolation and Intent-Driven Memory Management
by: Wu, Yiyang, et al.
Published: (2026)
by: Wu, Yiyang, et al.
Published: (2026)
PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
by: Song, Yixin, et al.
Published: (2023)
by: Song, Yixin, et al.
Published: (2023)
Guidelines for Building Indexes on Partially Cache-Coherent CXL Shared Memory
by: Wu, Fangnuo, et al.
Published: (2025)
by: Wu, Fangnuo, et al.
Published: (2025)
PhoenixOS: Concurrent OS-level GPU Checkpoint and Restore with Validated Speculation
by: Wei, Xingda, et al.
Published: (2024)
by: Wei, Xingda, et al.
Published: (2024)
Exploiting Dependency and Parallelism: Real-Time Scheduling and Analysis for GPU Tasks
by: Zhang, Yuanhai, et al.
Published: (2026)
by: Zhang, Yuanhai, et al.
Published: (2026)
Performance Isolation and Semantic Determinism in Efficient GPU Spatial Sharing
by: Yang, Zhenyuan, et al.
Published: (2026)
by: Yang, Zhenyuan, et al.
Published: (2026)
From Imperative to Declarative: Towards LLM-friendly OS Interfaces for Boosted Computer-Use Agents
by: Wang, Yuan, et al.
Published: (2025)
by: Wang, Yuan, et al.
Published: (2025)
HeteroPod: XPU-Accelerated Infrastructure Offloading for Commodity Cloud-Native Applications
by: Yang, Bicheng, et al.
Published: (2025)
by: Yang, Bicheng, et al.
Published: (2025)
Flare: Anomaly Diagnostics for Divergent LLM Training in GPU Clusters of Thousand-Plus Scale
by: Cui, Weihao, et al.
Published: (2025)
by: Cui, Weihao, et al.
Published: (2025)
Delegation with Trust<T>: A Scalable, Type- and Memory-Safe Alternative to Locks
by: Ahmad, Noaman, et al.
Published: (2024)
by: Ahmad, Noaman, et al.
Published: (2024)
Leveraging OS-Level Primitives for Robotic Action Management
by: Zheng, Wenxin, et al.
Published: (2025)
by: Zheng, Wenxin, et al.
Published: (2025)
VUDA: Breaking CUDA-Vulkan Isolation for Spatial Sharing of Compute and Graphics on the Same GPU
by: Xu, Bin, et al.
Published: (2026)
by: Xu, Bin, et al.
Published: (2026)
Towards High-Goodput LLM Serving with Prefill-decode Multiplexing
by: Chen, Yukang, et al.
Published: (2025)
by: Chen, Yukang, et al.
Published: (2025)
From Good to Great: Improving Memory Tiering Performance Through Parameter Tuning
by: Kanellis, Konstantinos, et al.
Published: (2025)
by: Kanellis, Konstantinos, et al.
Published: (2025)
LMetric: Simple is Better - Multiplication May Be All You Need for LLM Request Scheduling
by: Zhang, Dingyan, et al.
Published: (2026)
by: Zhang, Dingyan, et al.
Published: (2026)
CXLAimPod: CXL Memory is all you need in AI era
by: Yang, Yiwei, et al.
Published: (2025)
by: Yang, Yiwei, et al.
Published: (2025)
BLITZSCALE: Fast and Live Large Model Autoscaling with O(1) Host Caching
by: Zhang, Dingyan, et al.
Published: (2024)
by: Zhang, Dingyan, et al.
Published: (2024)
Towards Agentic OS: An LLM Agent Framework for Linux Schedulers
by: Zheng, Yusheng, et al.
Published: (2025)
by: Zheng, Yusheng, et al.
Published: (2025)
Taiji: A DPU Memory Elasticity Solution for In-production Cloud Environments
by: Zheng, Hao, et al.
Published: (2025)
by: Zheng, Hao, et al.
Published: (2025)
Nomad: Non-Exclusive Memory Tiering via Transactional Page Migration
by: Xiang, Lingfeng, et al.
Published: (2024)
by: Xiang, Lingfeng, et al.
Published: (2024)
GPUVM: GPU-driven Unified Virtual Memory
by: Nazaraliyev, Nurlan, et al.
Published: (2024)
by: Nazaraliyev, Nurlan, et al.
Published: (2024)
Age-Memory Trade-off in Read-Copy-Update
by: Ramani, Vishakha, et al.
Published: (2024)
by: Ramani, Vishakha, et al.
Published: (2024)
Trustworthy and Controllable Professional Knowledge Utilization in Large Language Models with TEE-GPU Execution
by: Cai, Yifeng, et al.
Published: (2025)
by: Cai, Yifeng, et al.
Published: (2025)
TierBPF: Page Migration Admission Control for Tiered Memory via eBPF
by: Wang, Xi, et al.
Published: (2026)
by: Wang, Xi, et al.
Published: (2026)
Sharpen the Spec, Cut the Code: A Case for Generative File System with SYSSPEC
by: Liu, Qingyuan, et al.
Published: (2025)
by: Liu, Qingyuan, et al.
Published: (2025)
Work-in-Progress: Multi-Deadline DAG Scheduling Model for Autonomous Driving Systems
by: Yano, Atsushi, et al.
Published: (2025)
by: Yano, Atsushi, et al.
Published: (2025)
AppFlow: Memory Scheduling for Cold Launch of Large Apps on Mobile and Vehicle Systems
by: Li, Xiaochen, et al.
Published: (2026)
by: Li, Xiaochen, et al.
Published: (2026)
SSV: Sparse Speculative Verification for Efficient LLM Inference
by: Wang, Zhibin, et al.
Published: (2026)
by: Wang, Zhibin, et al.
Published: (2026)
Assessing FIFO and Round Robin Scheduling:Effects on Data Pipeline Performance and Energy Usage
by: Choudhury, Malobika Roy, et al.
Published: (2024)
by: Choudhury, Malobika Roy, et al.
Published: (2024)
Vmem: A Lightweight Hot-Upgradable Memory Management for In-production Cloud Environment
by: Zheng, Hao, et al.
Published: (2025)
by: Zheng, Hao, et al.
Published: (2025)
Global Scheduling of Weakly-Hard Real-Time Tasks using Job-Level Priority Classes
by: Moyano, V. Gabriel, et al.
Published: (2024)
by: Moyano, V. Gabriel, et al.
Published: (2024)
CHRONOS: Compensating Hardware Related Overheads with Native Multi Timer Support for Real-Time Operating Systems
by: Heider, Kay, et al.
Published: (2025)
by: Heider, Kay, et al.
Published: (2025)
The First Principle of Big Memory Systems
by: Hua, Yu
Published: (2023)
by: Hua, Yu
Published: (2023)
UrgenGo: Urgency-Aware Transparent GPU Kernel Launching for Autonomous Driving
by: Zhu, Hanqi, et al.
Published: (2025)
by: Zhu, Hanqi, et al.
Published: (2025)
ARMS: Adaptive and Robust Memory Tiering System
by: Yadalam, Sujay, et al.
Published: (2025)
by: Yadalam, Sujay, et al.
Published: (2025)
Similar Items
-
Holistic Heterogeneous Scheduling for Autonomous Applications using Fine-grained, Multi-XPU Abstraction
by: Han, Mingcong, et al.
Published: (2025) -
Microsecond-scale Dynamic Validation of Idempotency for GPU Kernels
by: Han, Mingcong, et al.
Published: (2024) -
Characterizing Network Requirements for GPU API Remoting in AI Applications
by: Wang, Tianxia, et al.
Published: (2024) -
Towards Efficient and Practical GPU Multitasking in the Era of LLM
by: Xing, Jiarong, et al.
Published: (2025) -
Optimizing Tree-structure Indexes for CXL-based Heterogeneous Memory with SINLK
by: Zhao, Haoru, et al.
Published: (2025)