Performance Isolation and Semantic Determinism in Efficient GPU Spatial Sharing
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Zhenyuan, Zheng, Wenxin, Li, Mingyu, Chen, Haibo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VUDA: Breaking CUDA-Vulkan Isolation for Spatial Sharing of Compute and Graphics on the Same GPU
by: Xu, Bin, et al.
Published: (2026)
by: Xu, Bin, et al.
Published: (2026)
Peformance Isolation for Inference Processes in Edge GPU Systems
by: Martín, Juan José, et al.
Published: (2026)
by: Martín, Juan José, et al.
Published: (2026)
PhoenixOS: Concurrent OS-level GPU Checkpoint and Restore with Validated Speculation
by: Wei, Xingda, et al.
Published: (2024)
by: Wei, Xingda, et al.
Published: (2024)
Towards Efficient and Practical GPU Multitasking in the Era of LLM
by: Xing, Jiarong, et al.
Published: (2025)
by: Xing, Jiarong, et al.
Published: (2025)
NCCLbpf: Verified, Composable Policy Execution for GPU Collective Communication
by: Zheng, Yusheng
Published: (2026)
by: Zheng, Yusheng
Published: (2026)
GPUOS: A GPU Operating System Primitive for Transparent Operation Fusion
by: Yang, Yiwei, et al.
Published: (2026)
by: Yang, Yiwei, et al.
Published: (2026)
RAGDoll: Efficient Offloading-based Online RAG System on a Single GPU
by: Yu, Weiping, et al.
Published: (2025)
by: Yu, Weiping, et al.
Published: (2025)
BLITZSCALE: Fast and Live Large Model Autoscaling with O(1) Host Caching
by: Zhang, Dingyan, et al.
Published: (2024)
by: Zhang, Dingyan, et al.
Published: (2024)
"Range as a Key" is the Key! Fast and Compact Cloud Block Store Index with RASK
by: Zhao, Haoru, et al.
Published: (2026)
by: Zhao, Haoru, et al.
Published: (2026)
TrEnv-X: Transparently Share Serverless Execution Environments Across Different Functions and Nodes
by: Huang, Jialiang, et al.
Published: (2025)
by: Huang, Jialiang, et al.
Published: (2025)
Microsecond-scale Dynamic Validation of Idempotency for GPU Kernels
by: Han, Mingcong, et al.
Published: (2024)
by: Han, Mingcong, et al.
Published: (2024)
LEFT-RS: A Lock-Free Fault-Tolerant Resource Sharing Protocol for Multicore Real-Time Systems
by: Chen, Nan, et al.
Published: (2025)
by: Chen, Nan, et al.
Published: (2025)
Telepathic Datacenters: Fast RPCs using Shared CXL Memory
by: Mahar, Suyash, et al.
Published: (2024)
by: Mahar, Suyash, et al.
Published: (2024)
Mewz: Lightweight Execution Environment for WebAssembly with High Isolation and Portability using Unikernels
by: Ueda, Soichiro, et al.
Published: (2024)
by: Ueda, Soichiro, et al.
Published: (2024)
Optimizing Task Scheduling in Heterogeneous Computing Environments: A Comparative Analysis of CPU, GPU, and ASIC Platforms Using E2C Simulator
by: Mohammadjafari, Ali, et al.
Published: (2024)
by: Mohammadjafari, Ali, et al.
Published: (2024)
Formal Definitions and Performance Comparison of Consistency Models for Parallel File Systems
by: Wang, Chen, et al.
Published: (2024)
by: Wang, Chen, et al.
Published: (2024)
GPUVM: GPU-driven Unified Virtual Memory
by: Nazaraliyev, Nurlan, et al.
Published: (2024)
by: Nazaraliyev, Nurlan, et al.
Published: (2024)
Nixie: Efficient, Transparent Temporal Multiplexing for Consumer GPUs
by: Xu, Yechen, et al.
Published: (2026)
by: Xu, Yechen, et al.
Published: (2026)
TIDAL: Recovering Temporal Phase for Cloud Block Storage Placement from LLM-Derived Semantics
by: Tan, Difan, et al.
Published: (2026)
by: Tan, Difan, et al.
Published: (2026)
Evaluating Serverless Machine Learning Performance on Google Cloud Run
by: Khatiwada, Prerana, et al.
Published: (2024)
by: Khatiwada, Prerana, et al.
Published: (2024)
SwitchFS: Asynchronous Metadata Updates for Distributed Filesystems with In-Network Coordination
by: Xu, Jingwei, et al.
Published: (2024)
by: Xu, Jingwei, et al.
Published: (2024)
Fork, Explore, Commit: OS Primitives for Agentic Exploration
by: Wang, Cong, et al.
Published: (2026)
by: Wang, Cong, et al.
Published: (2026)
FALCON: Pinpointing and Mitigating Stragglers for Large-Scale Hybrid-Parallel Training
by: Wu, Tianyuan, et al.
Published: (2024)
by: Wu, Tianyuan, et al.
Published: (2024)
Agent Centric Operating System -- a Comprehensive Review and Outlook for Operating System
by: Jia, Shian, et al.
Published: (2024)
by: Jia, Shian, et al.
Published: (2024)
HybridTier: an Adaptive and Lightweight CXL-Memory Tiering System
by: Song, Kevin, et al.
Published: (2023)
by: Song, Kevin, et al.
Published: (2023)
DPC: A Distributed Page Cache over CXL
by: Bergman, Shai, et al.
Published: (2026)
by: Bergman, Shai, et al.
Published: (2026)
ContiguousKV: Accelerating LLM Prefill with Granularity-Aligned KV Cache Management
by: Zou, Jing, et al.
Published: (2026)
by: Zou, Jing, et al.
Published: (2026)
MegaTrain: Full Precision Training of 100B+ Parameter Large Language Models on a Single GPU
by: Yuan, Zhengqing, et al.
Published: (2026)
by: Yuan, Zhengqing, et al.
Published: (2026)
LMetric: Simple is Better - Multiplication May Be All You Need for LLM Request Scheduling
by: Zhang, Dingyan, et al.
Published: (2026)
by: Zhang, Dingyan, et al.
Published: (2026)
Unlocking True Elasticity for the Cloud-Native Era with Dandelion
by: Kuchler, Tom, et al.
Published: (2025)
by: Kuchler, Tom, et al.
Published: (2025)
CPU-Limits kill Performance: Time to rethink Resource Control
by: Shetty, Chirag, et al.
Published: (2025)
by: Shetty, Chirag, et al.
Published: (2025)
EdgeFlow: Fast Cold Starts for LLMs on Mobile Devices
by: Yan, Yongsheng, et al.
Published: (2026)
by: Yan, Yongsheng, et al.
Published: (2026)
A Periodic Space of Distributed Computing: Vision & Framework
by: Salehi, Mohsen Amini, et al.
Published: (2026)
by: Salehi, Mohsen Amini, et al.
Published: (2026)
Equilibria: Fair Multi-Tenant CXL Memory Tiering At Scale
by: Zhao, Kaiyang, et al.
Published: (2026)
by: Zhao, Kaiyang, et al.
Published: (2026)
CvxCluster: Solving Large, Complex, Granular Resource Allocation Problems 100-1000x Faster
by: Nnorom Jr, Obi, et al.
Published: (2026)
by: Nnorom Jr, Obi, et al.
Published: (2026)
Ensuring Data Freshness in Multi-Rate Task Chains Scheduling
by: Hoffmann, José Luis Conradi, et al.
Published: (2026)
by: Hoffmann, José Luis Conradi, et al.
Published: (2026)
Rethinking Inter-Process Communication with Memory Operation Offloading
by: Park, Misun, et al.
Published: (2026)
by: Park, Misun, et al.
Published: (2026)
Why iCloud Fails: The Category Mistake of Cloud Synchronization
by: Borrill, Paul
Published: (2026)
by: Borrill, Paul
Published: (2026)
Rethinking Thread Scheduling under Oversubscription: A User-Space Framework for Coordinating Multi-runtime and Multi-process Workloads
by: Roca, Aleix, et al.
Published: (2026)
by: Roca, Aleix, et al.
Published: (2026)
Characterizing Metastable Faults and Failures
by: Farahbakhsh, Ali, et al.
Published: (2026)
by: Farahbakhsh, Ali, et al.
Published: (2026)
Similar Items
-
VUDA: Breaking CUDA-Vulkan Isolation for Spatial Sharing of Compute and Graphics on the Same GPU
by: Xu, Bin, et al.
Published: (2026) -
Peformance Isolation for Inference Processes in Edge GPU Systems
by: Martín, Juan José, et al.
Published: (2026) -
PhoenixOS: Concurrent OS-level GPU Checkpoint and Restore with Validated Speculation
by: Wei, Xingda, et al.
Published: (2024) -
Towards Efficient and Practical GPU Multitasking in the Era of LLM
by: Xing, Jiarong, et al.
Published: (2025) -
NCCLbpf: Verified, Composable Policy Execution for GPU Collective Communication
by: Zheng, Yusheng
Published: (2026)